Quince logo
QuincePosted 1 month ago

Staff Engineer - ML Infra / MLOps

$285–$218,000 year

On-sitePalo Alto, California, United States

Full TimeSenior LevelSmallRetail E-commerce

Job Summary

Architect the end-to-end ML platform foundation, owning design for model training, serving, and feature pipelines while ensuring scalability and extensibility. Build the core developer experience for data scientists, enabling seamless movement from idea to production with minimal friction. Drive technical excellence by setting CI/CD standards, managing Infrastructure as Code, and implementing deployment strategies like blue-green and canary rollouts. Optimize compute performance through GPU utilization improvements and model batching to maximize cost efficiency across workloads. Ensure production reliability by architecting systems that handle traffic surges and seasonal spikes with robust monitoring and automated recovery. Mentor junior and mid-level engineers through design reviews and code reviews to elevate collective technical standards. Lead root-cause analyses for production failures, driving systemic fixes and modeling rigorous on-call discipline.

Required Qualifications

  • 8+ years of industry experience
  • at least 4+ years of focused, hands-on work in ML Infrastructure, MLOps, or large-scale Data Platform engineering
  • Proven track record of designing and building MLOps platforms that support the full model lifecycle — from data ingestion and distributed training to real-time inference and model governance
  • Deep expertise in cloud-native infrastructure (preferably AWS)
  • Kubernetes (EKS)
  • Docker
  • Infrastructure as Code tools (Terraform/Pulumi)
  • Hands-on mastery of ML frameworks such as PyTorch, TensorFlow, Kubeflow, or SageMaker
  • Expertise in building Feature Stores and high-throughput data pipelines (Spark, Flink, Kafka)
  • strong understanding of training/serving skew and data consistency
  • Expert-level knowledge of CI/CD for ML, including model versioning, experiment tracking, and deployment strategies such as blue-green and canary rollouts
  • Demonstrated ability to optimize GPU utilization, implement model batching, and systematically reduce cloud infrastructure costs
  • Strong operational instincts, with a history of improving reliability through rigorous on-call practices, proactive monitoring, and root-cause analysis
  • You understand the hustle of a startup and are good at handling ambiguity
  • You are a curious, quick learner who loves to experiment and thrives at a rapid pace
  • Employment is contingent upon successful completion of a background check

Desired Qualifications

  • preferably AWS

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce