Site Reliability Engineer III
On-sitePlano, Texas, United States
Job Summary
Guide peers in adopting site reliability engineering best practices and design level structures. Collaborate with software engineers to design, develop, and implement deployment approaches using automated continuous integration and delivery pipelines. Configure, maintain, and optimize applications and infrastructure through code, including infrastructure and network as code. Resolve complex problems proactively using service level indicators and objectives, while applying enterprise-authorized AI capabilities to accelerate incident triage, troubleshooting, and post-incident analysis. Identify patterns in operational signals to prioritize reuse-first improvements tied to SLO outcomes and iteratively improve availability, reliability, and scalability. This role supports the IP team at JPMorgan Chase, focusing on modernizing mission-critical systems through cloud infrastructure and data management platforms.
Required Qualifications
- Formal training or certification on site reliability engineering concepts
- 3+ years applied experience
- Exposure to or hands-on experience in supporting SRE practices for Data management/migration platforms and products
- Familiarity in on-prem/public cloud infrastructure components such as Compute(Linux/windows), Storage, Networks and Database
- Understanding of how to apply SRE fundamentals — including monitoring, incident response, capacity awareness, and toil identification
- Ability to define and track relevant SLOs/SLIs
- Proficient in site reliability culture and principles
- Familiarity with how to implement site reliability within an application or platform
- Proficient in at least one programming language or configuration/resource management tools such as Python, Ansible and Terraform
- Experience in observability including white and black box monitoring, service level objective alerting, and telemetry collection using tools such as Grafana, Dynatrace, Prometheus, Datadog, Splunk, etc.
- Experience with platforms and applications hosted on public/private/hybrid cloud environments
- Experience with container orchestration technologies such as Kubernetes, ECS, and Docker
- Experience with continuous integration and continuous delivery tools such as Jenkins, GitLab, or Terraform
- Familiarity with troubleshooting common networking technologies and issues
- Experience with continuous integration and continuous delivery tooling
- Familiarity with container and container orchestration
- Familiarity with troubleshooting common networking technologies and issues
Desired Qualifications
- Exposure to or hands-on experience in supporting SRE practices for Data management/migration platforms and products
- Familiarity in on-prem/public cloud infrastructure components such as Compute(Linux/windows), Storage, Networks and Database
- Understanding of how to apply SRE fundamentals — including monitoring, incident response, capacity awareness, and toil identification
- Ability to define and track relevant SLOs/SLIs
- Proficient in site reliability culture and principles
- Familiarity with how to implement site reliability within an application or platform
- Proficient in at least one programming language or configuration/resource management tools such as Python, Ansible and Terraform
- Experience in observability including white and black box monitoring, service level objective alerting, and telemetry collection using tools such as Grafana, Dynatrace, Prometheus, Datadog, Splunk, etc.
- Experience with platforms and applications hosted on public/private/hybrid cloud environments
- Experience with container orchestration technologies such as Kubernetes, ECS, and Docker
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.