Quadric logo
QuadricPosted 2 weeks ago

AI Performance Modeling Engineer

$150,000–$200,000 year

On-siteBurlingame, California, United States

Full TimeSmall

Job Summary

Build analytical, cycle-level Python models of AI inference workloads executing on next-generation GPNPU hardware before silicon exists. Derive hardware lane bindings for compute, memory bandwidth, and interconnects from first principles, while modeling tensor placement, tiling, and data movement across memory tiers. Incorporate architectural details for vision networks and LLMs, including operator mix, sparsity, and quantization formats. Calibrate performance models against instruction-set simulators and profiling traces to meet accuracy targets, then write and defend technical studies that directly inform architecture and product decisions. Balance single-stream latency against scaled throughput performance. Own a full workload's model end-to-end within 6–12 months, ensuring predictions stay within 10–15% of actual measurements.

Required Qualifications

  • Strong Python skills with experience writing, validating, and calibrating numerical or quantitative models in code
  • Solid grasp of memory hierarchies, bandwidth/latency trade-offs, pipelining, and execution bottlenecks
  • Comfort writing clear technical studies that state and defend evidence-based conclusions
  • BS, MS, or Ph.D. in Computer Science, Electrical Engineering, Computer Engineering, or equivalent practical experience

Desired Qualifications

  • Prior experience with GPUs, custom AI accelerators, CUDA, or Triton kernels
  • Familiarity with roofline analysis, back-of-the-envelope estimation, or architecture simulators (e.g., gem5, Timeloop, MAESTRO, Accel-Sim)
  • Background in compiler internals (cost models, autotuners) or proficiency in C++
  • Published performance studies or technical write-ups

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce