AI Performance Modeling Engineer
$150,000–$200,000 year
On-siteBurlingame, California, United States
Job Summary
Build analytical, cycle-level Python models of AI inference workloads executing on next-generation GPNPU hardware before silicon exists. Derive hardware lane bindings for compute, memory bandwidth, and interconnects from first principles, while modeling tensor placement, tiling, and data movement across memory tiers. Incorporate architectural details for vision networks and LLMs, including operator mix, sparsity, and quantization formats. Calibrate performance models against instruction-set simulators and profiling traces to meet accuracy targets, then write and defend technical studies that directly inform architecture and product decisions. Balance single-stream latency against scaled throughput performance. Own a full workload's model end-to-end within 6–12 months, ensuring predictions stay within 10–15% of actual measurements.
Required Qualifications
- Strong Python skills with experience writing, validating, and calibrating numerical or quantitative models in code
- Solid grasp of memory hierarchies, bandwidth/latency trade-offs, pipelining, and execution bottlenecks
- Comfort writing clear technical studies that state and defend evidence-based conclusions
- BS, MS, or Ph.D. in Computer Science, Electrical Engineering, Computer Engineering, or equivalent practical experience
Desired Qualifications
- Prior experience with GPUs, custom AI accelerators, CUDA, or Triton kernels
- Familiarity with roofline analysis, back-of-the-envelope estimation, or architecture simulators (e.g., gem5, Timeloop, MAESTRO, Accel-Sim)
- Background in compiler internals (cost models, autotuners) or proficiency in C++
- Published performance studies or technical write-ups
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.