Harell Data logo
Harell DataPosted 3 weeks ago

Software Engineer, AI Infrastructure

On-sitePalo Alto, California, United States

Full TimeSmall

Job Summary

Orchestrate GPU workloads on Kubernetes, managing resource allocation, scheduling, and multi-tenancy for the compute layer. Build the inference layer by handling model loading, autoscaling, and batching to ensure low latency and high throughput. Own the end-to-end ML pipeline, including data ingestion, preprocessing, training, and recovery from multi-node job failures. Debug failing fine-tuning jobs and develop real-time observability for model performance and resource health. Manage incident response and on-call rotation to maintain platform reliability as usage scales. Lead build-vs-buy decisions on infrastructure and security while setting engineering standards. This early engineer role reports directly to the CTO, focusing on turning scattered scientific data into domain-specific foundation models for Harell Data's managed platform.

Required Qualifications

  • 5+ years building and operating production infrastructure, with a focus on ML workloads: training, inference, or data pipelines
  • Hands-on experience with Kubernetes on AWS or GCP, ideally with GPU workloads
  • Strong CS fundamentals and system design chops
  • Comfortable with ambiguity — you've worked somewhere where the playbook didn't exist yet
  • Location: Palo Alto, CA

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce