Senior Machine Learning Engineer, Infrastructure
$212,000–$318,000 year
HybridNew York City, New York, United States or San Francisco, California, United States
Job Summary
Architect, scale, and maintain high-throughput, low-latency live inference infrastructure to support fan discovery and content ranking systems. Own the end-to-end feature store lifecycle from ingestion to production serving while designing observability frameworks to detect performance gaps and latency spikes. Collaborate with product, data engineering, and trust & safety teams to translate requirements into scalable infrastructure solutions and automate model deployment with reliability testing. Debug complex relevance systems when monitoring identifies bottlenecks, ensuring high availability and consistency between online and offline features. Work within the Relevance team to build robust, maintainable code in Python and improve developer velocity through shared infrastructure and clear documentation.
Required Qualifications
- deep experience building, deploying, and maintaining production-grade ML infrastructure at scale
- specifically with low-latency live inference pipelines
- feature store architectures
- strong background in distributed systems
- backend engineering
- ability to write robust, maintainable code in Python
- systematic approach to debugging complex, high-throughput systems
- ability to debug performance bottlenecks
- reliability issues
- in-office 3 days per week
- San Francisco or New York
Desired Qualifications
- energetic by building '0 to 1' infrastructure systems
- strong communication skills
- effective at creating clear documentation for system architectures
- infrastructure strategies
- growth mindset
- keen eye for detail in code reviews
- passion for empowering your teammates
- improving developer velocity
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.