AI and HPC Systems Performance Engineer
HybridBengaluru, Karnataka, India
Job Summary
Install, configure, and optimize complex AI infrastructure components including GPU servers, storage systems, high-speed networking, and AI software stacks. Develop automation scripts, deployment frameworks, and Infrastructure-as-Code solutions to streamline AI platform provisioning and workload execution. Perform system-level performance characterization and optimization of AI training and inference workloads on HPE platforms utilizing GPU accelerators and distributed computing technologies. Design, execute, and analyze performance benchmarks for AI/ML workloads, including large language models, multimodal models, and Retrieval-Augmented Generation pipelines. Capture, analyze, and interpret system telemetry, performance metrics, logs, traces, and profiling data to identify bottlenecks and optimization opportunities. Collaborate with customers, partners, and internal engineering teams to characterize, troubleshoot, and optimize AI solutions deployed on HPE infrastructure. Author technical reports, white papers, reference architectures, and benchmark studies to guide performance optimization and solution design.
Required Qualifications
- Hybrid work arrangement with average 2 days per week from an HPE office
- Senior or principal-level engineer
- 8+ years of experience
- Strong experience with Linux system administration and command-line environments across multiple enterprise Linux distributions
- Experience with modern AI/ML frameworks and ecosystems including PyTorch, JAX, Hugging Face Transformers, and related technologies
- Experience with AI model training, inference, benchmarking, performance characterization, and optimization
- Experience with data analysis, statistical methods, experiment design, and performance modeling techniques
- Experience conducting technical research and evaluating emerging AI technologies, frameworks, and hardware platforms
- Experience with high-performance networking technologies including InfiniBand, RDMA, RoCE, and Mellanox/NVIDIA networking solutions
- Strong analytical, troubleshooting, and root-cause analysis skills
- Proficiency in one or more programming or scripting languages such as Python, Bash, Go, C++, or similar
- Experience working with complex, distributed, multi-layer software systems and AI infrastructure stacks
- Experience using performance profiling, tracing, observability, and benchmarking tools to analyze system and application performance
- Experience with containerized and orchestrated environments including Docker, Kubernetes, and related cloud-native technologies
- Experience analyzing and optimizing AI workloads running on GPU-accelerated systems
- Experience with distributed training and inference frameworks and large-scale AI/LLM workloads
- Experience with GPU accelerator technologies, memory hierarchies, and AI software stacks including CUDA, NCCL, and related ecosystem tools
- Experience with distributed GPU environments and multi-node AI clusters
- Experience with LLM serving frameworks such as vLLM, TensorRT-LLM, SGLang, or similar technologies
- Aptitude for self-learning; Learns new concepts quickly
- Excellent written and verbal communication; mastery in English
- Ability to work well in a team environment and perform well under pressure
- MS/ME/MTech or PhD in Computer Science, Computer Engineering, Electrical Engineering, Data Science, Artificial Intelligence, or a related technical discipline
- 5+ years of experience in AI/ML infrastructure, performance engineering, high-performance computing (HPC), or related technical fields
- Experience with large-scale AI training and inference environments supporting foundation models and large language models (LLMs)
- Experience with HPE platforms, AI Factory architectures, or enterprise AI infrastructure solutions
- Experience with parallel and distributed storage technologies, including Weka, Lustre, BeeGFS, GPFS, or similar high-performance file systems
- Experience with AI benchmarking methodologies and industry benchmarks such as MLPerf
- Experience developing reference architectures, technical papers, benchmark studies, or performance guidance documentation
- Experience working directly with customers, partners, and cross-functional engineering teams in highly collaborative environments
- Experience optimizing performance across multi-GPU and multi-node AI environments using InfiniBand, RDMA, GPUDirect, or equivalent technologies
- Ability to work independently in globally distributed teams with minimal supervision
Desired Qualifications
- Experience troubleshooting complex, multi-tier software systems and distributed AI environments
- Experience with one or more of PyTorch, JAX, Hugging Face Transformers, DeepSpeed, Megatron-LM, Ray, vLLM, SGLang, TensorRT-LLM, Dynamo, ONNX Runtime, Kubernetes, Redis, Vector Databases, Retrieval-Augmented Generation (RAG) architectures, distributed training and inference, and large-scale AI/LLM workloads
- Experience characterizing and optimizing performance across multi-GPU and distributed AI environments utilizing modern GPU interconnect, networking, and storage technologies
- Experience with tuning system performance in a benchmarking environment
- Experience with distributed GPU environments and multi-node AI clusters
- Experience with LLM serving frameworks such as vLLM, TensorRT-LLM, SGLang, or similar technologies
- Experience with parallel and distributed storage technologies, including Weka, Lustre, BeeGFS, GPFS, or similar high-performance file systems
- Experience with AI benchmarking methodologies and industry benchmarks such as MLPerf
- Experience developing reference architectures, technical papers, benchmark studies, or performance guidance documentation
- Experience working directly with customers, partners, and cross-functional engineering teams in highly collaborative environments
- Experience optimizing performance across multi-GPU and multi-node AI environments using InfiniBand, RDMA, GPUDirect, or equivalent technologies
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.