Applied AI Engineer, Kernel Performance
$150,000–$225,000 year
On-siteSan Jose, California, United States
Job Summary
Own the system that converts new model architectures into verified, production-ready kernels optimized for Etched hardware. Build autonomous agents that understand proprietary hardware, design experiments, generate implementations, compile and profile them, and diagnose bottlenecks to reach peak performance faster than traditional workflows. Design evals covering correctness, numerical stability, latency, and efficiency while curating datasets from complete optimization trajectories. Ship model-generated improvements to production and quantify their impact on end-to-end system performance. Partner with architecture teams to shape hardware-software roadmaps and continuously evaluate new model releases.
Required Qualifications
- A track record of solving hard problems across stacks and domains
- Comfort with both Python and low-level code
- Kernel experience
- Fluency using AI to learn and ramp on new problems
- Moving fluidly between research exploration, agentic experimentation, low-level debugging, and production execution
- Must be available for in-person work in San Jose (Santana Row)
Desired Qualifications
- First principles thinking on accelerator performance
- Hands-on experience building and shipping LLM-based agents or AI tooling that real users depend on in production environments
- An eval-driven mindset
- Fine-tuning or post-training
- RAG over proprietary data
- Multi-agent orchestration
- High agency and comfort with ambiguity
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.