CB Smart Recruit logo
CB Smart RecruitPosted 1 month ago

Senior / Lead / Principal Platform Engineer (DevOps / Cloud Infrastructure)

$200,000–$300,000 year

On-siteLos Angeles, California, United States

Full TimeSenior LevelSmall

Job Summary

Design, build, and operate scalable cloud infrastructure supporting production AI/ML workloads, including Kubernetes architecture, networking, security, and upgrades. Own the internal developer platform to improve engineering productivity and deployment velocity through self-service automation and Infrastructure-as-Code solutions using Terraform. Implement modern CI/CD and GitOps workflows while optimizing GPU provisioning and resource management for model training and inference. Lead incident investigation, root-cause analysis, and observability implementation to drive operational reliability across distributed systems. Partner with AI engineers and technical leadership to champion automation, scalability, and security best practices.

Required Qualifications

  • Bachelor's degree in Computer Science, Software Engineering, Information Technology, or a related technical discipline (Master's preferred)
  • 5+ years of experience building and operating production cloud infrastructure, Platform Engineering, DevOps, or Site Reliability Engineering (SRE) environments
  • Strong software engineering foundation with experience building automation, tooling, services, or developer platforms using Python, Go, Bash, or similar languages
  • Demonstrated ownership of production Kubernetes clusters, including architecture, networking, upgrades, scaling, and operational support
  • Hands-on experience designing and building Infrastructure-as-Code solutions using Terraform, including authoring reusable modules
  • Strong experience designing and building CI/CD and GitOps pipelines—not simply maintaining existing pipelines
  • Deep experience with Google Cloud Platform (GCP) and/or AWS
  • Strong understanding of containerization technologies including Docker and Kubernetes
  • Experience building and operating production-scale distributed systems
  • Strong troubleshooting skills across cloud infrastructure, Kubernetes, networking, and applications
  • Experience with observability platforms such as Prometheus, Grafana, Datadog, ELK, or equivalent
  • Excellent communication and collaboration skills
  • Applicants must be legally authorized to work in the United States
  • Visa sponsorship is not available for this role
  • On-site (5 days per week)
  • West Hollywood / Los Angeles, CA

Desired Qualifications

  • Master's degree (preferred)
  • AI/ML infrastructure and GPU-accelerated workloads
  • NVIDIA GPU infrastructure and CUDA environments
  • Internal developer platforms and self-service infrastructure
  • GitOps methodologies
  • AI-native development tools such as Claude Code, Cursor, GitHub Copilot, or Codex
  • Security-focused environments including DevSecOps practices
  • Air-gapped, sovereign, or highly regulated deployment environments
  • Defense, aerospace, government, or other mission-critical industries
  • FedRAMP, ITAR, CMMC, or similar compliance frameworks
  • Serverless architectures and distributed systems

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce