Senior Software Engineer, Machine Learning Services
On-siteLondon, England, United Kingdom
Job Summary
Design, build, and operate the core Machine Learning Services platform, including Rust-based API gateways, Python compute workers, and distributed job queues. Solve hard concurrency and performance problems to ensure bulletproof production workloads for high-volume inference and unattended model training. Develop custom storage abstractions over cloud object stores and enhance asynchronous job-queueing systems using compare-and-swap primitives. Work directly with product and ML science teams to build scalable infrastructure for GenAI models and specialized classifiers. Dive deep into the stack from Kubernetes and gRPC to ONNX-based inference on GPU-accelerated hardware. Write clean, efficient, and rigorously tested code with a focus on simplicity and correctness.
Required Qualifications
- 5+ years of engineering and architecting large-scale, distributed commercial services
- Deep proficiency in a systems-level language (Rust, C++, Go)
- A willingness and curiosity to become an expert in Rust
- Strong Python skills
- Real-world experience with cloud ecosystems (Azure, AWS, or GCP)
- Experience with containerization (Docker, Kubernetes)
- A firm grasp of concurrency, multithreading, and asynchronous programming
- A pragmatic understanding of computer science fundamentals
- The ability to articulate an opinion on what makes good code and good architecture
- The ability to challenge assumptions and contribute to a culture of continuous improvement
Desired Qualifications
- Experience with MLOps, particularly the challenges of managing the lifecycle of models in a multi-tenant, high-availability system
- Familiarity with building ML inference services, model serialization (e.g., ONNX), and GPU programming (CUDA)
- Experience with building or working on custom storage or job-queueing systems
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.