Lambda logo
LambdaPosted 1 week ago

Staff Software Engineer - Compute

$314,000–$465,000 year

HybridSan Francisco, California, United States or San Jose, California, United States

Full TimeSenior LevelSmall

Job Summary

Design and implement a highly available GPU and CPU host and instance lifecycle control plane, bridging distributed systems with semiconductor architecture to enable seamless cloud provisioning. Guide technical decisions on BIOS/Firmware, system boot methodologies, and DPU utilization to optimize host performance and reliability. Lead the development of a multi-tenant security model and provide mentorship to senior engineers across teams while collaborating with product and data center organizations to translate customer requirements into scalable infrastructure capabilities. Set engineering standards and lead design reviews for mission-critical cloud software at scale. Requires 10+ years of experience in compute control plane distributed systems and proficiency in C/C++, Rust, Python, or Go. This role supports Lambda's mission to make compute as ubiquitous as electricity for AI researchers and enterprises.

Required Qualifications

  • 10+ years of experience working on compute control plane distributed systems used for deploying and lifecycle managing heterogeneous compute platforms into data-centers, built for resilience at scale.
  • Deep expertise in durable execution models and distributed systems used in cloud-service provisioning.
  • Basic knowledge of software defined networking fundamentals that informs secure, multi-tenant distributed systems.
  • Proven track record of leading large-scale semi-conductor hardware enablement and deployment initiatives.
  • Proven experience in deploying net-new data-centers into a global compute platform (not just working in existing data-centers).
  • Proficiency in one of more of the following programming languages: C/C++, Rust, Python, Go.
  • Presence in our Bellevue, San Francisco, or San Jose office location 4 days per week
  • Lambda's designated work from home day is currently Tuesday

Desired Qualifications

  • Knowledge of Nvidia's AI Factory architectural components (including GPU hosts, CPU hosts, SuperNICs (ConnectX and Bluefield DPUs , and switches).
  • Knowledge of Nvidia's AI Factory software offerings (like DOCA, DOCA SNAP, CUDA, et al.).
  • Knowledge of Linux kernel internals, device drivers, and virtualization technologies (KVM, QEMU), kernel bypass technologies (like SR-IOV, DPDK, SPDK).
  • Experience with Cloud Service Provider Kubernetes offerings.
  • Knowledge of high-performance networking (InfiniBand, RoCE) and storage protocols (NVMe-oF).

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce