Senior Site Reliability Engineer, AI Platform
HybridBengaluru, Karnataka, India or Hyderabad, Telangana, India
Job Summary
Own and drive long-term infrastructure strategy for the AI Platform team, designing scalable systems and standards that reduce operational burden. Improve operational posture of cloud-native services including containerized workloads, serverless functions, and managed ML infrastructure by instrumenting them with observability tooling covering metrics, logs, and distributed tracing. Establish SLAs and error budgets while leading incident response, post-mortems, and systemic reliability improvements. Collaborate with platform engineering to harden infrastructure against security findings and maintain compliance. Design, build, and maintain CI/CD pipelines for AI Platform services and SDKs, then automate toil for infrastructure provisioning and deployment operations. Mentor junior team members to raise the engineering bar across DevOps practices.
Required Qualifications
- Strong communication skills and the ability to work across engineering, security, and product teams and distill unclear requirements into clear engineering goals
- 5+ years of experience in cloud and infrastructure engineering roles
- Strong proficiency in Python and shell scripting, with hands-on experience building and deploying automation tools that eliminate toil and streamline DevOps workflows
- Infrastructure-as-code experience with CloudFormation, Terraform, ARM or SAM
- Deep hands-on experience with cloud infrastructure
- Solid experience with CI/CD tooling and infrastructure deployment automation
- Familiar with containerization (Docker, Kubernetes) and cloud-native deployment patterns
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.