Software Engineer/Senior Software Engineer, Data & ML Platform
On-siteSanta Clara, California, United States
Job Summary
Manage production Kubernetes clusters by handling node lifecycles, upgrades, and troubleshooting while implementing GitOps and infrastructure-as-code practices. Drive end-to-end ownership of GPU and ML workload scheduling, including queueing, priorities, fractional sharing, and autoscaling across multi-tenant environments. Leverage hands-on experience with Ray, Kubeflow, Delta Lake, Apache Iceberg, Apache Spark, and Argo Workflows to optimize large-scale distributed data-processing systems. This role supports the Data & ML Platform team in building robust, scalable infrastructure for advanced analytics and machine learning workloads.
Required Qualifications
- BS, MS, or PhD in Computer Science or a related technical field, or equivalent practical experience
- Hands-on experience operating production Kubernetes clusters — node lifecycle, upgrades, troubleshooting — plus GitOps and infrastructure-as-code experience
- Experience with GPU or ML workload scheduling, queueing and priorities, fractional GPU sharing, autoscaling, or multi-tenant resource management
- Self-driven with a strong sense of ownership: a quick learner who is eager to take responsibility and drive projects forward end to end
Desired Qualifications
- Experience with Ray or Kubeflow
- Experience with lakehouse technologies such as Delta Lake or Apache Iceberg
- Experience operating large-scale distributed data-processing and workflow systems, with hands-on depth in a system such as Apache Spark and working knowledge of Argo Workflows or an equivalent orchestrator
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.