Software Engineer, AI Infrastructure
On-sitePalo Alto, California, United States
Job Summary
Orchestrate GPU workloads on Kubernetes, managing resource allocation, scheduling, and multi-tenancy for the compute layer. Build the inference layer by handling model loading, autoscaling, and batching to ensure low latency and high throughput. Own the end-to-end ML pipeline, including data ingestion, preprocessing, training, and recovery from multi-node job failures. Debug failing fine-tuning jobs and develop real-time observability for model performance and resource health. Manage incident response and on-call rotation to maintain platform reliability as usage scales. Lead build-vs-buy decisions on infrastructure and security while setting engineering standards. This early engineer role reports directly to the CTO, focusing on turning scattered scientific data into domain-specific foundation models for Harell Data's managed platform.
Required Qualifications
- 5+ years building and operating production infrastructure, with a focus on ML workloads: training, inference, or data pipelines
- Hands-on experience with Kubernetes on AWS or GCP, ideally with GPU workloads
- Strong CS fundamentals and system design chops
- Comfortable with ambiguity — you've worked somewhere where the playbook didn't exist yet
- Location: Palo Alto, CA
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.