Senior / Lead / Principal Platform Engineer (DevOps / Cloud Infrastructure)
$200,000–$300,000 year
On-siteLos Angeles, California, United States
Job Summary
Design, build, and operate scalable cloud infrastructure supporting production AI/ML workloads, including Kubernetes architecture, networking, security, and upgrades. Own the internal developer platform to improve engineering productivity and deployment velocity through self-service automation and Infrastructure-as-Code solutions using Terraform. Implement modern CI/CD and GitOps workflows while optimizing GPU provisioning and resource management for model training and inference. Lead incident investigation, root-cause analysis, and observability implementation to drive operational reliability across distributed systems. Partner with AI engineers and technical leadership to champion automation, scalability, and security best practices.
Required Qualifications
- Bachelor's degree in Computer Science, Software Engineering, Information Technology, or a related technical discipline (Master's preferred)
- 5+ years of experience building and operating production cloud infrastructure, Platform Engineering, DevOps, or Site Reliability Engineering (SRE) environments
- Strong software engineering foundation with experience building automation, tooling, services, or developer platforms using Python, Go, Bash, or similar languages
- Demonstrated ownership of production Kubernetes clusters, including architecture, networking, upgrades, scaling, and operational support
- Hands-on experience designing and building Infrastructure-as-Code solutions using Terraform, including authoring reusable modules
- Strong experience designing and building CI/CD and GitOps pipelines—not simply maintaining existing pipelines
- Deep experience with Google Cloud Platform (GCP) and/or AWS
- Strong understanding of containerization technologies including Docker and Kubernetes
- Experience building and operating production-scale distributed systems
- Strong troubleshooting skills across cloud infrastructure, Kubernetes, networking, and applications
- Experience with observability platforms such as Prometheus, Grafana, Datadog, ELK, or equivalent
- Excellent communication and collaboration skills
- Applicants must be legally authorized to work in the United States
- Visa sponsorship is not available for this role
- On-site (5 days per week)
- West Hollywood / Los Angeles, CA
Desired Qualifications
- Master's degree (preferred)
- AI/ML infrastructure and GPU-accelerated workloads
- NVIDIA GPU infrastructure and CUDA environments
- Internal developer platforms and self-service infrastructure
- GitOps methodologies
- AI-native development tools such as Claude Code, Cursor, GitHub Copilot, or Codex
- Security-focused environments including DevSecOps practices
- Air-gapped, sovereign, or highly regulated deployment environments
- Defense, aerospace, government, or other mission-critical industries
- FedRAMP, ITAR, CMMC, or similar compliance frameworks
- Serverless architectures and distributed systems
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.