Bank of America logo
Bank of AmericaPosted 1 month ago

Senior Site Reliability Engineer

$152,600–$191,500 year

On-siteCharlotte, North Carolina, United States or Jersey City, New Jersey, United States

Full TimeSenior LevelLarge

Job Summary

Design and implement advanced reliability patterns for Azure landing zones, private networking, DNS, firewalls, and regional resiliency. Lead complex platform initiatives including secondary-region readiness, egress/ingress observability, and enterprise dashboard automation. Define SLIs, SLOs, alerting standards, and service health reporting while developing reusable Terraform modules and automation frameworks. Drive deep technical investigations for major incidents and partner with security teams to integrate IAM, policy-as-code, and audit readiness into operations. Mentor SRE engineers, create executive-ready technical summaries, and establish standards for runbooks and post-incident reviews. Operate as a senior technical authority for GCP platform reliability, resiliency, and operational excellence, collaborating with development and infrastructure teams to prioritize reliability stories and reduce manual support effort.

Required Qualifications

  • 4+ years of experience in cloud infrastructure engineering, platform engineering, or cloud operations, with exposure to Google Cloud Platform (GCP)
  • Strong hands-on experience with Infrastructure as Code (IaC), including practical use of Terraform or Terraform Enterprise for infrastructure provisioning
  • Solid understanding of software engineering fundamentals, including version control, code quality, and basic testing practices for infrastructure code
  • Experience developing and maintaining Terraform modules and infrastructure configurations to support automated cloud environments
  • Familiarity with CI/CD pipelines for infrastructure deployment, including automated build, test, and release processes
  • Working knowledge of DevSecOps practices, including integrating security and compliance checks into automated workflows
  • Good understanding of GCP services and cloud architecture fundamentals, including networking (VPCs, IAM, load balancing)
  • Exposure to policy-as-code, governance, and compliance requirements in enterprise environments
  • Experience supporting automation and standardization efforts to improve consistency and efficiency in cloud deployments
  • Understanding of monitoring, logging, and observability tools to support system reliability and performance
  • Hands-on experience with incident response, troubleshooting, and root cause analysis in cloud or distributed systems

Desired Qualifications

  • Ability to collaborate effectively with engineering, architecture, and security teams to support reliable and secure platform operations
  • Strong problem-solving and analytical skills, with the ability to diagnose and resolve infrastructure issues
  • Effective communication skills, with the ability to work within cross-functional teams and document technical solutions clearly
  • Interest in learning and applying emerging technologies and automation techniques (including AI/ML where applicable) to improve platform reliability

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce