Senior Site Reliability Engineer
HybridBengaluru, Karnataka, India
Bengaluru, Karnataka, IndiaHybridFull TimeSenior LevelMasters DegreeTECHLarge
Full TimeSenior LevelMasters DegreeLargeTECH
Job Summary
Lead and onboard services to reliability tenets while establishing Service Level Objectives and Agreements. Design scalable, secure infrastructure using cloud-native best practices and automate operational tasks with Python and Gen AI tooling. Collaborate with development teams to ensure systems are resilient and performant, implementing monitoring, alerting, and tracing solutions. Lead incident response efforts, conduct root cause analysis, and participate in 24/7 on-call rotation for production support.
Required Qualifications
- BS/MS in Computer Science or Equivalent
- 6+ years of experience in Site Reliability Engineering, managing infrastructure and services at scale
- History of end-to-end project delivery
- Experience managing Hadoop and Kubernetes infrastructure and related services, or equivalent experience
- Advanced knowledge of Linux, Networking, and Containers
- Proficiency in at least one high-level programming language (Python, GoLang etc.)
- Strong troubleshooting and debugging skills
- Fluency in English and excellent communication skills
- An understanding or hands-on experience with Prompt engineering in software development
- Familiarity with AI Native IDEs or AI Assistants such as Cursor, GitHub CoPilot, and Claude
- Ability to establish, monitor, and improve Service Level Objectives (SLOs), Indicators (SLIs), and Agreements (SLAs) in line with business needs
- Containerization technologies and orchestration platforms, mainly Kubernetes and EKS (CKA, CKAD, CKS certifications are valued)
- Experience with automation and Infrastructure as Code (IaC) tools, such as AWS CloudFormation, Terraform, Puppet, Chef, Spacelift, etc
- Experience with Python, Go, Bash/Shell scripting, or other automation tools/languages
- Familiarity with AWS services like EC2, RDS, ELB, CloudFront, Lambda, etc
- Proficiency in monitoring and troubleshooting complex distributed systems
- Experience with Grafana, ELK stack, Prometheus, or others
- Strong understanding of designing resilient and fault-tolerant systems
- Expertise in debugging complex distributed systems
Desired Qualifications
- Experience in any of the following is valued, but not fully required: Ability to establish, monitor, and improve Service Level Objectives (SLOs), Indicators (SLIs), and Agreements (SLAs) in line with business needs
- Experience in any of the following is valued, but not fully required: Containerization technologies and orchestration platforms, mainly Kubernetes and EKS (CKA, CKAD, CKS certifications are valued)
- Experience in any of the following is valued, but not fully required: Experience with automation and Infrastructure as Code (IaC) tools, such as AWS CloudFormation, Terraform, Puppet, Chef, Spacelift, etc
- Experience in any of the following is valued, but not fully required: Experience with Python, Go, Bash/Shell scripting, or other automation tools/languages
- Experience in any of the following is valued, but not fully required: Familiarity with AWS services like EC2, RDS, ELB, CloudFront, Lambda, etc
- Experience in any of the following is valued, but not fully required: Proficiency in monitoring and troubleshooting complex distributed systems
- Experience in any of the following is valued, but not fully required: Experience with Grafana, ELK stack, Prometheus, or others
- Experience in any of the following is valued, but not fully required: Strong understanding of designing resilient and fault-tolerant systems
- Experience in any of the following is valued, but not fully required: Expertise in debugging complex distributed systems
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.