Member of Technical Staff, Kernel Engineering
$148,000–$296,000 year
On-siteSingapore, Singapore
Singapore, SingaporeOn-siteFull Time$148,000–$296,000 yearMid LevelBachelors DegreeStartup
Full TimeMid LevelBachelors DegreeStartup
Job Summary
Performance engineer to squeeze every FLOP out of modern accelerators by writing kernels and low-level optimizations for vLLM; work across GPUs and emerging silicon, collaborating with hardware teams to maximize performance. Requires CUDA kernel experience, strong GPU architecture knowledge, proficient C++ and Python, and experience with profiling tools and benchmarking. Visa sponsorship is offered case-by-case; role is based in Singapore with in-person operations.
Required Qualifications
- Bachelor's degree or equivalent experience in computer science, engineering, or similar
- Deep experience writing CUDA kernels or equivalent (CuTeDSL, Triton, TileLang, Pallas)
- Strong understanding of GPU architecture: memory hierarchy, warp scheduling, tiling, tensor cores
- Proficiency in C++ and Python with demonstrated ability to write high-performance code
- Experience with profiling tools (Nsight, rocprof) and performance optimization methodologies
- Obsession with benchmarks and squeezing every percentage point of speedup
Desired Qualifications
- Deep experience writing CUDA kernels or equivalent (CuTeDSL, Triton, TileLang, Pallas)
- Strong understanding of GPU architecture: memory hierarchy, warp scheduling, tiling, tensor cores
- Proficiency in C++ and Python with demonstrated ability to write high-performance code
- Experience with profiling tools (Nsight, rocprof) and performance optimization methodologies
- Obsession with benchmarks and squeezing every percentage point of speedup
- Experience with ML-specific kernel optimization (FlashAttention, fused kernels)
- Knowledge of quantization techniques (INT8, FP8, mixed-precision)
- Familiarity with multiple accelerator platforms (NVIDIA, AMD, TPU, Intel)
- Experience with compiler technologies (LLVM, MLIR, XLA)
- Kernel-related contributions to vLLM or other inference engine projects
- Contributions to open-source GPU, ML systems, or compiler optimization projects
- Written deep technical blogs on GPU optimization
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.