Delan Associates logo
Delan AssociatesPosted 3 weeks ago

LLM Inference & GPU Systems Consultant

On-siteCharlotte, North Carolina, United States

ContractSmall

Job Summary

Optimize NVIDIA H200 GPU clusters for token generation pipelines, managing prefill/decode efficiency and KV cache strategies. Deploy and maintain inference engines including vLLM and TensorRT-LLM while tuning throughput, batching, and latency. Operate the OpenShift AI ecosystem as the primary container platform for GenAI workloads and orchestrate GPU resources using RunAI and Kubernetes. Oversee the complete Hugging Face model lifecycle from onboarding to retirement within this enterprise private environment. This 6-month onsite role in Charlotte, NC requires 3+ days per week and focuses exclusively on inferencing infrastructure.

Required Qualifications

  • Must be onsite at client in Charlotte, NC at least 3 days/week
  • 8+ years experience working as an LLM Systems Engineer or AI Infrastructure Runtime Engineer
  • 8+ years hands-on experience with NVIDIA H200 clusters and runtime optimization techniques (KV Cache, prefill/decode)
  • Proficiency in OpenShift AI and GPU orchestration tools like RunAI
  • Strong experience with modern inference frameworks, specifically vLLM and TensorRT-LLM
  • Proven track record managing the Hugging Face deployment lifecycle

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce