Evals Infrastructure Tech Lead / Manager
$500,000–$850,000 year
HybridSan Francisco, California, United States
Job Summary
Design and implement high-performance data processing infrastructure for large language model training, including core primitives like tokenization and deduplication. Build robust systems for data quality assurance, monitoring, and distributed computing across web-scale datasets. Collaborate with research teams to optimize architectures for model performance and ensure reproducibility and traceability in data preparation. Provide front-line leadership by managing day-to-day execution, prioritizing work in a dynamic environment, and coaching reports on professional growth. Maintain deep technical understanding of AI safety implications and the team's stack to drive targeted contributions and system reliability.
Required Qualifications
- 1+ years of management experience in a technical environment, particularly performance or distributed systems
- Strong software engineering skills with experience in building distributed systems
- Expertise in Python and Rust
- Deep understanding of cloud computing platforms and distributed systems architecture
- Experience with high-throughput, fault-tolerant system design
- Strong background in performance optimization and system scaling
- Excellent problem-solving skills and attention to detail
- Strong communication skills and ability to work in a collaborative environment
- Experience with language model training infrastructure
- Strong background in distributed systems and parallel computing
- Expertise in tokenization algorithms and techniques
- Experience building high-throughput, fault-tolerant systems
- Deep knowledge of monitoring and observability practices
- Experience with infrastructure-as-code and configuration management
- Bachelor's degree or an equivalent combination of education, training, and/or experience
- A field relevant to the role as demonstrated through coursework, training, or professional experience
Desired Qualifications
- Have significant experience building and maintaining large-scale distributed systems
- Are passionate about system reliability and performance
- Enjoy solving complex technical challenges at scale
- Are comfortable working with ambiguous requirements and evolving specifications
- Take ownership of problems and drive solutions independently
- Are excited about contributing to the development of safe and ethical AI systems
- Can balance technical excellence with practical delivery
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.