AI Cloud Senior DevOps Engineer
RemoteUnited States or Singapore
Job Summary
Design, implement, and maintain end-to-end CI/CD pipelines for software applications and machine learning models, automating build, test, deployment, and rollback processes. Build, optimize, and scale cloud-native infrastructure using Kubernetes and Docker to manage specialized computing resources like GPU clusters for AI workloads. Take ownership of high-availability design, disaster recovery strategies, and capacity planning while championing Infrastructure as Code practices with Terraform, Ansible, and Helm. Architect comprehensive monitoring, logging, and alerting systems to provide deep visibility into system health and AI model metrics. Act as the technical lead during complex incidents, spearheading rapid troubleshooting, root cause analysis, and preventative remediation plans. Work closely with R&D, Data Science, and Security teams to streamline workflows and establish robust security standards including Zero Trust access controls and compliance with SOC2 and ISO27001 frameworks.
Required Qualifications
- Bachelor's degree or above in Computer Science, Engineering, or a related technical field
- 5+ years of hands-on experience in DevOps, Site Reliability Engineering (SRE), or Cloud Infrastructure roles
- Expert-level knowledge of Linux operating systems
- Expert-level knowledge of core networking principles (TCP/IP, DNS, HTTP, Load Balancing, VPCs)
- Deep mastery of Docker
- Deep mastery of Kubernetes orchestration
- Proven proficiency in designing and managing infrastructure on major Public or Hybrid Cloud platforms (e.g., AWS, GCP, Azure, Alibaba Cloud)
- Strong coding and scripting capabilities in at least one major language (Go, Python, Shell, etc.)
- Systematic and practical understanding of CI/CD methodologies
- Systematic and practical understanding of Infrastructure as Code (IaC)
- Systematic and practical understanding of Observability paradigms
- Systematic and practical understanding of Site Reliability Engineering (SRE) principles
- Exceptional problem-solving abilities
- Sharp technical judgment
- Excellent cross-team communication skills
Desired Qualifications
- Familiarity with MLOps practices
- Experience with model serving/inferencing frameworks (e.g., vLLM, TGI, Triton Inference Server)
- Experience managing GPU clusters for AI/ML workloads
- Proven track record working with large-scale distributed systems or high-concurrency environments (e.g., Fintech, Trading, Real-time processing, or AI platforms)
- Hands-on experience in designing and building Internal Developer Platforms (IDP)
- Deep familiarity with Zero Trust architecture
- Experience with automated security testing (DevSecOps)
- Experience implementing strict compliance frameworks (e.g., SOC2, ISO27001)
- Prior experience acting as a Technical Lead
- Experience mentoring junior engineers
- Experience managing DevOps teams
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.