Lambda logo
LambdaPosted 1 week ago

Senior Site Reliability Engineer

$240,000–$356,000 year

HybridSan Francisco, California, United States or Bellevue, Washington, United States

Full TimeSenior LevelSmall

Job Summary

Build and operate monitoring and alerting for cluster health across fabric, GPU, power, and job-level signals to detect and respond to issues proactively. Remotely deploy and configure large-scale HPC clusters for AI workloads using automation, while automating cluster lifecycle management for operating systems, firmware, drivers, and networking via Ansible and Terraform. Create runbooks and automated remediations for common failure modes, then troubleshoot and resolve cluster issues involving InfiniBand, RoCE, NCCL, and GPU-direct environments. Participate in on-call rotations and lead incident response for cluster-level problems, contributing to Standard Operating Procedures and feeding requirements back to engineering teams.

Required Qualifications

  • 7+ years of experience in Site Reliability Engineering, HPC Engineering, DevOps, or a similar role
  • Have a strong understanding of modern AI infrastructure, from GPU architectures to hardware performance optimization
  • Strong understanding of Linux-based systems in a distributed environment
  • Are experienced configuring and troubleshooting InfiniBand (IB), RoCE, CLOS fabrics, 100GbE, Ethernet/switching, GPU-direct, and NCCL environments
  • Solid understanding of Python and Go, with experience working with SWE teams to improve internal tooling
  • Experience with monitoring and alerting tools (e.g., Prometheus, Grafana, Clickhouse)
  • Proficiency in automation and configuration management tools (e.g., Ansible, Terraform)
  • Have excellent problem-solving and troubleshooting skills and an innate attention to detail
  • Passion for continuous improvement and innovation
  • Note: This position requires presence in our San Francisco or Bellevue office location 4 days per week; Lambda's designated work from home day is currently Tuesday

Desired Qualifications

  • Experience with machine learning / deep learning frameworks (PyTorch, TensorFlow) and benchmarking tools (DeepSpeed, MLPerf)
  • Knowledge of containerization and orchestration technologies (e.g., Docker, Kubernetes)
  • Experience building and/or operating HPC resources
  • Depth in the NVIDIA hardware and firmware ecosystem
  • Experience with data center power and thermal design
  • Background in chaos engineering or similar reliability testing methodologies
  • Understanding of compliance frameworks (SOC 2, ISO 27001, etc.)

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce