Zscaler logo
ZscalerPosted 3 weeks ago

Sr. Staff Software Development Engineer - AI Platform

HybridBengaluru, Karnataka, India

Full TimeSenior LevelLargeCybersecurity Services

Job Summary

Design, build, and maintain scalable, secure AWS infrastructure for AI/ML workloads using Terraform, managing EKS, Lambda, ECS, VPC, and IAM resources. Own GitLab CI/CD pipelines for automated builds, security scanning, and multi-environment deployments while architecting centralized observability stacks with Prometheus and Grafana. Lead incident response, root-cause analysis, and postmortems to improve MTTA/MTTR, and define DORA metrics to drive delivery improvements. Implement self-healing, auto-scaling infrastructure for LLM inference and vector databases, partnering with AI/ML teams on production-grade RAG deployments. Mentor engineers through design reviews and code audits to uphold platform governance standards.

Required Qualifications

  • 8+ years of experience as a Platform Engineer / Site Reliability Engineer / DevOps Engineer
  • 3+ years supporting AI/ML or data platform infrastructure
  • deep hands-on expertise in AWS (EKS, Lambda, ECS, VPC, IAM, S3)
  • Strong hands-on experience with Infrastructure as Code using Terraform (modules, workspaces, remote state)
  • designing/maintaining CI/CD pipelines with GitLab CI/CD (or equivalent), including runners and deployment automation
  • Proven expertise in centralized observability using Prometheus and Grafana (custom exporters, dashboard design, alerting rules)
  • experience designing and tuning alerting/on-call systems (Alertmanager, PagerDuty, OpsGenie)
  • owning incident response for production services
  • Solid, hands-on understanding of DevOps/DORA metrics (deployment frequency, lead time for changes, change failure rate, MTTR)
  • how to use them to drive engineering improvement
  • Experience deploying and operating containerized applications on Kubernetes (EKS)
  • including Helm charts and autoscaling (HPA, Cluster Autoscaler, Karpenter)
  • strong scripting skills in Python (and/or Go/Bash)
  • the ability to work independently and lead complex infrastructure initiatives
  • excellent communication skills befitting a Staff-level engineer

Desired Qualifications

  • Demonstrated curiosity and active exploration of AI tools
  • a proven history of integrating new technologies to enhance daily workflows and augment problem-solving
  • Experience with LiteLLM (including multi-geo/multi-region deployment patterns)
  • distributed compute frameworks like Ray.io/Anyscale
  • workflow engines like Temporal
  • agentic frameworks such as LangGraph, LangChain, or LlamaIndex
  • Working knowledge of vector database platforms (Qdrant, Pinecone, Weaviate)
  • memory/context layers like ZEP
  • LLMOps/MLOps frameworks and practices (model registry, prompt versioning, feedback loops, MLflow/Kubeflow)
  • LLM observability/evaluation tooling such as Arize Phoenix
  • Exposure to GitOps workflows (ArgoCD/Flux)
  • log aggregation and distributed tracing (Loki, ELK/OpenSearch, Jaeger/Tempo)
  • cloud cost optimization (FinOps)
  • security/compliance frameworks (SOC 2, ISO 27001) in cloud-native environments

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce