DigitalOcean logo
DigitalOceanPosted 3 weeks ago

Senior Forward Deployed Engineer I (AI Inference)

HybridBengaluru, Karnataka, India

Full TimeSenior LevelLargeCloud Computing

Job Summary

Architect and deploy production-grade, multi-tenant LLM inference engines using Kubernetes-native frameworks like llm-d, NVIDIA Dynamo, Ray Serve, and vLLM. Lead forward-deployed engagements by embedding with external tech leads to debug latency spikes, profile GPU memory utilization, and refactor inference code for high-concurrency workloads. Optimize cluster-scale serving through prefill/decode disaggregation, KV-cache-aware routing, and quantization techniques to maximize tokens-per-second per dollar. Translate customer edge cases into reusable internal blueprints and contribute performance fixes to open-source inference ecosystems. Travel up to 30% for customer workshops and strategic collaborations while maintaining availability until noon Eastern Time.

Required Qualifications

  • 6+ years in AI/ML systems
  • Deep understanding of why cluster-scale serving is hard (e.g., partitioning KV-cache across workers, fast cross-pod KV transfer, and inference-aware load balancing)
  • Hands-on experience with vLLM, llm-d, SGLang, TensorRT-LLM, or Modular MAX
  • Solid grasp of internals like continuous batching and paged attention
  • Expert-level proficiency in Python or GoLang
  • Familiarity with gRPC
  • Experience running critical services on Kubernetes in high-scale environments
  • Ability to travel up to 30% for customer engagements, strategic workshops, conferences, and internal collaboration
  • Ability to consistently overlap with North American business hours, including availability until at least noon Eastern Time

Desired Qualifications

  • 6+ years of experience working in Forward Deployed Engineering, AI Inference architect, Technical Consulting roles supporting production AI systems
  • Ability to translate complex business tasks into AI engineering solutions and collaborate directly with client teams (CTOs, AI Leads)
  • Preference for delivering production-ready code, low-latency container images, and deployment blueprints over slide decks
  • Comfortable navigating fast-moving environments and tuning model workloads for diverse accelerator architectures
  • Experience collaborating with GPU vendors, infrastructure providers, model vendors, or ecosystem partners on benchmarking, optimization, technical validation, or launch readiness initiatives

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce