Patterson-UTI logo
Patterson-UTIPosted 1 week ago

Senior Site Reliability Engineer NEX

On-siteHouston, Texas, United States

Full TimeSenior LevelBachelors DegreeLargeOil and Gas Services

Job Summary

Design, implement, and operate scalable, resilient systems on Google Cloud Platform while defining service-level indicators, objectives, and error budgets. Perform capacity planning, disaster recovery design, and secure cloud networking implementation to optimize resource utilization and cost efficiency. Build infrastructure using Terraform, automate operational activities, and improve CI/CD pipelines for reliable software delivery. Develop actionable alerts, maintain dashboards, and participate in a sustainable on-call rotation to respond to production incidents and lead postmortems. Partner with software engineering, security, and product teams to establish reliability standards and promote shared responsibility for production operations.

Required Qualifications

  • Three or more years of experience in Site Reliability Engineering, platform engineering, DevOps, cloud engineering, production software engineering, or a similar role.
  • Experience operating highly available systems in a 24/7 production environment.
  • Hands-on experience operating workloads on Google Cloud Platform or another major public cloud platform.
  • Strong experience managing compute, networking and data GCP services workloads
  • Strong experience with containerization and orchestration technologies, including Docker and Kubernetes.
  • Experience building and managing infrastructure with Terraform or a comparable infrastructure-as-code tool.
  • Proficiency in Python, Go, Java, or another comparable programming language.
  • Experience implementing or operating CI/CD pipelines using GitHub Actions, Azure DevOps, Bitbucket Pipelines, or comparable tools.
  • Experience implementing observability using metrics, logs, traces, dashboards, and alerts.
  • Experience participating in on-call rotations, responding to incidents, and contributing to postmortems.
  • Understanding of SLIs, SLOs, error budgets, and other SRE principles.
  • Ability to troubleshoot complex issues across application, infrastructure, network, data, and cloud-service layers.
  • Ability to communicate effectively with engineering teams, business stakeholders, and operational personnel.
  • Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field, or equivalent practical experience.
  • 3+ years of experience in Site Reliability Engineering, platform engineering, cloud engineering, or DevOps.
  • 3+ years of experience operating production workloads in GCP.
  • Ability to understand and communicate in English at a level sufficient to issue, receive, and respond to safety-related and operations-related instructions.

Desired Qualifications

  • Google Cloud and/or Kubernetes certifications.
  • Experience supporting data-intensive, streaming, analytics, or event-driven platforms.
  • Experience establishing production-readiness, incident-management, or reliability-review processes.
  • Experience working in the energy, oil and gas, industrial, IoT, field operations, or other operationally critical industries.
  • Experience supporting technology environments that integrate cloud platforms with remote sites, field equipment, industrial systems, or edge computing.

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce