Sr. Staff Software Development Engineer - AI Platform
HybridBengaluru, Karnataka, India
Job Summary
Design, build, and maintain scalable, secure AWS infrastructure for AI/ML workloads using Terraform, managing EKS, Lambda, ECS, VPC, and IAM resources. Own GitLab CI/CD pipelines for automated builds, security scanning, and multi-environment deployments while architecting centralized observability stacks with Prometheus and Grafana. Lead incident response, root-cause analysis, and postmortems to improve MTTA/MTTR, and define DORA metrics to drive delivery improvements. Implement self-healing, auto-scaling infrastructure for LLM inference and vector databases, partnering with AI/ML teams on production-grade RAG deployments. Mentor engineers through design reviews and code audits to uphold platform governance standards.
Required Qualifications
- 8+ years of experience as a Platform Engineer / Site Reliability Engineer / DevOps Engineer
- 3+ years supporting AI/ML or data platform infrastructure
- deep hands-on expertise in AWS (EKS, Lambda, ECS, VPC, IAM, S3)
- Strong hands-on experience with Infrastructure as Code using Terraform (modules, workspaces, remote state)
- designing/maintaining CI/CD pipelines with GitLab CI/CD (or equivalent), including runners and deployment automation
- Proven expertise in centralized observability using Prometheus and Grafana (custom exporters, dashboard design, alerting rules)
- experience designing and tuning alerting/on-call systems (Alertmanager, PagerDuty, OpsGenie)
- owning incident response for production services
- Solid, hands-on understanding of DevOps/DORA metrics (deployment frequency, lead time for changes, change failure rate, MTTR)
- how to use them to drive engineering improvement
- Experience deploying and operating containerized applications on Kubernetes (EKS)
- including Helm charts and autoscaling (HPA, Cluster Autoscaler, Karpenter)
- strong scripting skills in Python (and/or Go/Bash)
- the ability to work independently and lead complex infrastructure initiatives
- excellent communication skills befitting a Staff-level engineer
Desired Qualifications
- Demonstrated curiosity and active exploration of AI tools
- a proven history of integrating new technologies to enhance daily workflows and augment problem-solving
- Experience with LiteLLM (including multi-geo/multi-region deployment patterns)
- distributed compute frameworks like Ray.io/Anyscale
- workflow engines like Temporal
- agentic frameworks such as LangGraph, LangChain, or LlamaIndex
- Working knowledge of vector database platforms (Qdrant, Pinecone, Weaviate)
- memory/context layers like ZEP
- LLMOps/MLOps frameworks and practices (model registry, prompt versioning, feedback loops, MLflow/Kubeflow)
- LLM observability/evaluation tooling such as Arize Phoenix
- Exposure to GitOps workflows (ArgoCD/Flux)
- log aggregation and distributed tracing (Loki, ELK/OpenSearch, Jaeger/Tempo)
- cloud cost optimization (FinOps)
- security/compliance frameworks (SOC 2, ISO 27001) in cloud-native environments
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.