Staff AI Systems Engineer
$172,800–$216,000 year
On-siteSan Jose, California, United States
Job Summary
Architect, deploy, and manage resilient infrastructure services for large-scale AI model training and low-latency inference. Optimize hardware utilization across multi-cloud environments using specialized frameworks and advanced LLM serving engines. Maintain end-to-end tooling including MLflow to streamline the AI development lifecycle while partnering with researchers to productionize cutting-edge models. Debug complex performance bottlenecks at the hardware-software interface and establish monitoring systems for high-throughput data pipelines. This role supports Archer's mission to build an all-electric vertical takeoff and landing aircraft, bridging the gap between raw data and high-performance AI models.
Required Qualifications
- BS/MS/PhD degree in Computer Science, Software Engineering, or a related field
- 5+ years of professional software engineering experience with a dedicated focus on AI/ML systems, high-performance computing (HPC), or ML infrastructure
- Familiarity with hyper-scaler infrastructure (AWS)
- Familiarity with specialized AI-centric bare-metal and GPU clouds (Nebius AI Cloud)
- Hands-on experience with containerization (Docker)
- Production-grade orchestration (Kubernetes)
- Cloud-agnostic cluster abstractors like SkyPilot
- Deep architectural understanding of large language models
- Experience building high-throughput data pipelines to support large-scale training
- Proficiency in SQL
- Proficiency in NoSQL
- Proficiency in columnar storage formats optimized for ML (e.g., Parquet)
Desired Qualifications
- Familiarity with audio processing
- Familiarity with speech-to-text frameworks
- Familiarity with Automatic Speech Recognition (ASR) pipelines
- Prior experience or a deep technical interest in aerospace, aviation, or autonomous systems (e.g., safety-critical software, edge-AI deployments)
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.