Bright Vision Technologies logo
Bright Vision TechnologiesPosted 1 week ago
EXPIRED

ML Platform Engineer

$100,000–$160,000 year

RemoteUnited States

Full TimeMasters DegreeSmall

Job Summary

Design and operate high-performance, highly reliable inference platforms for serving large machine learning models in production. Optimize inference performance using continuous batching, paged attention, speculative decoding, and request multiplexing while implementing multi-tenant routing, rate limiting, and quality-of-service policies. Build autoscaling and capacity management systems that balance latency, throughput, and cost, and tune GPU utilization, memory management, and KV cache strategies for LLM serving workloads. Integrate model serving with API gateways, identity systems, and observability platforms, and drive end-to-end observability including latency histograms, queue dynamics, GPU utilization, and error tracking. Develop deployment workflows including canary releases, shadow testing, and automated rollback, and operate incident response for high-availability AI services. Collaborate with ML and product teams to support new model releases and capability rollouts, and implement security controls including request signing, content filtering, and abuse detection at the serving layer.

Required Qualifications

  • Bachelor's or Master's degree in Computer Science or a related field
  • 10 or more years of experience in distributed systems, infrastructure, or ML platform engineering
  • Strong proficiency in Python and a systems language such as Go, Rust, or C++
  • Deep experience operating high-throughput, low-latency services in production
  • Hands-on experience with LLM or large model inference frameworks such as vcLLM or TensorRT-LLM
  • Strong understanding of GPU architecture, memory hierarchies, and accelerator utilization
  • Familiarity with Kubernetes, autoscaling, and modern cloud platforms
  • Experience with observability stacks including metrics, tracing, and structured logging
  • Solid grounding in performance engineering and capacity planning
  • Strong communication and incident response skills

Desired Qualifications

  • Open-source contributions to model serving infrastructure
  • Experience with multi-region or globally distributed AI serving
  • Familiarity with model quantization, distillation, and compression techniques
  • Exposure to FinOps for AI workloads and cost-efficient serving design
  • Experience supporting external-facing AI APIs at scale

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Find similar roles