Designworks Talent logo
Designworks TalentPosted 1 month ago

Inference Engineer

$150,000–$250,000 year

HybridBellevue, Washington, United States

Full TimeSenior LevelStartup

Job Summary

Build and operate production-grade model-serving systems supporting high-throughput, low-latency AI workloads. Optimize inference infrastructure for token throughput, latency, scalability, and cost efficiency across different model architectures. Design systems that maximize GPU utilization while maintaining predictable performance and reliability. Improve platform scalability and operational maturity as customer demand grows. Partner with AI training, GPU performance, and orchestration teams to ensure smooth transitions from model development to production serving. Develop monitoring, alerting, and operational practices to maintain reliable inference services. Investigate and resolve performance, reliability, and capacity challenges across inference workloads. Contribute to architecture decisions and engineering standards as the platform evolves.

Required Qualifications

  • Experience building and operating production machine learning inference or model-serving systems at scale
  • Strong understanding of the performance trade-offs involved in serving large AI models, including latency, throughput, memory utilization, and cost efficiency
  • Experience designing reliable distributed systems or production infrastructure
  • Understanding of GPU-backed AI workloads and the challenges of scaling inference systems
  • Strong engineering fundamentals and the ability to independently own complex technical problems
  • Comfortable working in a fast-moving environment where systems and processes are being built from the ground up
  • U.S. work authorization
  • Hybrid role based in the Bellevue, WA area
  • Approximately three days per week in the office
  • Candidates elsewhere in the U.S. who are open to relocation

Desired Qualifications

  • Experience with modern inference-serving frameworks such as vLLM, TensorRT-LLM, Triton Inference Server, or similar technologies
  • Experience optimizing LLM inference workloads or large-scale AI serving platforms
  • Background operating API-based AI products or high-volume production services
  • Experience with GPU scheduling, distributed systems, Kubernetes, or cloud infrastructure platforms
  • Familiarity with model optimization techniques such as quantization, batching, caching, or performance tuning
  • Experience working at a hyperscaler, AI lab, GPU cloud provider, or large-scale ML infrastructure organization

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce