LLM Inference & GPU Systems Consultant
On-siteCharlotte, North Carolina, United States
Job Summary
Optimize NVIDIA H200 GPU clusters for token generation pipelines, managing prefill/decode efficiency and KV cache strategies. Deploy and maintain inference engines including vLLM and TensorRT-LLM while tuning throughput, batching, and latency. Operate the OpenShift AI ecosystem as the primary container platform for GenAI workloads and orchestrate GPU resources using RunAI and Kubernetes. Oversee the complete Hugging Face model lifecycle from onboarding to retirement within this enterprise private environment. This 6-month onsite role in Charlotte, NC requires 3+ days per week and focuses exclusively on inferencing infrastructure.
Required Qualifications
- Must be onsite at client in Charlotte, NC at least 3 days/week
- 8+ years experience working as an LLM Systems Engineer or AI Infrastructure Runtime Engineer
- 8+ years hands-on experience with NVIDIA H200 clusters and runtime optimization techniques (KV Cache, prefill/decode)
- Proficiency in OpenShift AI and GPU orchestration tools like RunAI
- Strong experience with modern inference frameworks, specifically vLLM and TensorRT-LLM
- Proven track record managing the Hugging Face deployment lifecycle
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.