Senior Site Reliability Engineer
$152,600–$191,500 year
On-siteJersey City, New Jersey, United States or Plano, Texas, United States
Job Summary
Design solutions to visualize production support metrics and develop software processes to remediate toil identified by collaborating with key partners. Partner with Development and Infrastructure teams to create error budget policies, prioritize reliability stories below SLO thresholds, and suggest code optimizations for service visibility. Identify capacity bottlenecks, vulnerabilities, and reliability improvement opportunities while assessing monitoring for new changes and enhancing application system monitoring designs. Engage as a subject matter expert in incident triage, failure scenario modeling, and root cause diagnosis for complex, high-impact investigations. Lead complex platform reliability initiatives including secondary-region readiness, egress/ingress observability, and enterprise dashboard automation. Define and mature SLIs, SLOs, reliability indicators, and alerting standards for Azure platform services. Develop reusable Terraform modules, automation frameworks, and CI/CD patterns that improve consistency, compliance, and operational quality. Drive observability improvements using Azure Monitor, Log Analytics, Dynatrace, and Resource Graph. Identify systemic reliability risks and translate them into engineering roadmaps, remediation plans, and operational controls. Partner with security and governance teams to integrate IAM, policy-as-code, vulnerability remediation, and audit readiness into Azure platform operations. Mentor SRE engineers and raise the technical bar for automation, troubleshooting, documentation, resiliency design, and production support. Create executive-ready technical summaries, reliability narratives, and recommendations for leadership review.
Required Qualifications
- Advanced experience in Azure platform engineering, SRE, cloud infrastructure, or enterprise cloud operations
- Deep knowledge of Microsoft Azure architecture, including networking, identity, compute, PaaS, monitoring, security, governance, and resiliency patterns
- Strong experience designing and developing Terraform modules and infrastructure-as-code automation in enterprise environments
- Strong understanding of Azure landing zones, hub-and-spoke networking, ExpressRoute or enterprise connectivity, private endpoints, DNS, routing, firewalls, and workload isolation
- Advanced observability experience with Azure Monitor, Log Analytics, Dynatrace, dashboards, alerting, metrics, and platform telemetry
- Experience with SRE operating models, SLIs, SLOs, incident response, problem management, toil reduction, and production-readiness reviews
- Strong scripting or programming experience with Python, PowerShell, Bash, Java, or similar languages
- Experience operating highly available Azure IaaS and PaaS services in enterprise-scale environments
- Ability to influence architecture and engineering decisions across multiple technical teams
- Strong communication skills with the ability to translate complex engineering topics into actionable recommendations
- 1st shift (United States of America)
- 40 hours per week
Desired Qualifications
- Microsoft Azure certification strongly preferred
- Experience in financial services, regulated technology, or large enterprise infrastructure organizations
- Experience with AKS, ACR, Kubernetes, container networking, CI/CD pipelines, and container observability
- Experience with Azure AI Foundry, OpenAI/GenAI platform operations, model-serving observability, or AI platform readiness
- Experience with FinOps, governance dashboards, resource hygiene, cost visibility, security posture, and compliance reporting
- Experience creating technical roadmaps, reliability scorecards, production-readiness frameworks, or executive-level reliability reporting
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.