Lambda logo
LambdaPosted 2 weeks ago

Senior Software Engineer - Storage Control Plane

$266,000–$395,000 year

HybridSan Francisco, California, United States or San Jose, California, United States

Full TimeSenior LevelSmall

Job Summary

Design and build a vendor-agnostic control plane that provisions, scales, heals, and meters storage across platforms like VAST Data, WEKA, and Ceph. Define an internal abstraction layer hiding vendor-specific APIs behind a declarative interface, then build reconciliation-loop orchestration using Kubernetes controllers to manage capacity, tenancy, and placement across data centers. Own multi-tenant isolation end-to-end, including namespace partitioning, per-tenant QoS, and noisy-neighbor detection. Instrument the fleet with SLI/SLO definitions and performance regression detection to make petabyte-scale operations debuggable. This role requires presence in San Francisco, San Jose, or Bellevue office 4 days per week, with Tuesday as the designated work-from-home day.

Required Qualifications

  • Bachelor's or Master's degree in Computer Science or a related field
  • 5+ years of experience in software development for storage systems
  • Proven experience with distributed systems programming and concepts such as load balancers, data-durability, consensus algorithms, fault tolerance, and data consistency
  • Strong programming skills in languages such as C, C++, Go, or Python
  • Experience with Linux kernel internals and system-level programming
  • Experience with one or more storage protocols (e.g. S3, NFS) and file systems such as Ceph, DAOS, or similar
  • Familiarity with containerization technologies like Docker and Kubernetes and running production workloads in these environments
  • Familiarity with CI/CD and QA practices for distributed systems development environments
  • Note: This position requires presence in our San Francisco, San Jose, or Bellevue office location 4 days per week; Lambda's designated work-from-home day is currently Tuesday

Desired Qualifications

  • Experience with AI/ML workloads and the unique storage challenges they present
  • Knowledge of data center networking and high-speed interconnects (e.g., InfiniBand, RoCE)
  • Experience with performance tuning and optimization of storage systems
  • Familiarity with hardware acceleration technologies, specifically GPUs and DPUs
  • Production experience with VAST Data, WEKA, DDN, Pure, NetApp, or IBM Storage Scale
  • Ceph at 100 PB+ in HPC or AI environments
  • CXL memory pooling, computational storage, ZNS SSDs, EDSFF
  • Published or presented at SNIA SDC, FAST, USENIX ATC, LSFMM+BPF, OCP, SC, or similar

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce