ML Infrastructure Engineer
Remote
Job Summary
Own inference and model-serving infrastructure end to end, from initial design through production deployment and ongoing scaling. Build systems that enable AI agents to run reliably and efficiently under high concurrency, identifying and resolving bottlenecks in collaboration with ML and platform teams. Optimize production environments for latency, throughput, and reliability while driving observability, monitoring, and debugging practices across the stack. This full-time, on-site role in San Mateo, CA requires 5+ years of hands-on experience with tools like Triton, TorchServe, or KServe, and proficiency in Python, Go, or Rust. You will shape foundational infrastructure decisions for a seed-stage enterprise AI company serving critical industries like banking and healthcare.
Required Qualifications
- 5+ years of hands-on experience building and operating machine learning inference systems, model-serving platforms, or ML infrastructure in production environments
- Demonstrated experience designing and scaling inference-serving infrastructure using tools such as TensorFlow Serving, TorchServe, Triton, KServe, or equivalent custom systems
- Proven ability to optimize production ML systems for latency, throughput, and reliability at scale
- Experience with containerization and orchestration (Docker, Kubernetes) for deploying and scaling ML workloads
- Background in distributed systems that handle high concurrency and dynamic resource allocation under load
- Proficiency with monitoring and observability tooling — e.g., Prometheus, Grafana, ELK stack, distributed tracing
- Experience deploying and managing ML systems on cloud platforms (AWS, GCP, or Azure)
- Proficiency in at least one systems or backend language: Python, Go, Rust, C++, or Java
- This is a full-time, on-site role based in San Mateo, CA
- On-site collaboration is an important part of how this small, fast-moving team operates
- Visa sponsorship is not available for this position
Desired Qualifications
- Experience with knowledge graphs, semantic search, or graph databases (e.g., Neo4j, Amazon Neptune, or similar)
- Background in real-time inference or low-latency serving requirements
- Familiarity with agentic AI systems, autonomous agents, or multi-step reasoning pipelines
- Experience with enterprise data infrastructure, data pipelines, or data integration platforms
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.