CoreWeave logo
CoreWeavePosted 4 weeks ago

Senior Systems Engineer, Test Frameworks & Validation Platform

$153,000–$204,000 year

On-siteLivingston, New Jersey, United States

Part TimeSenior LevelLargeTECH

Job Summary

Build and operate the test framework that qualifies host images on real physical hardware before they reach the GPU fleet. Own the in-house, Kubernetes-native system that pins to a host, boots configurations, runs containerized checks, and records structured results. Extend coverage into HPC verification, Slurm-on-Kubernetes, and firmware validation while hardening the core to eliminate flaky tests. Manage the end-to-end results path, including dashboards and triage surfaces for rapid root cause analysis. Explore AI-native testing with LLM-driven log triage and failure classification. Collaborate with firmware, kernel, and HPC teams to embed testing into their release processes.

Required Qualifications

  • 3+ years of experience building test infrastructure, systems software, or platform tooling at scale, with real ownership of frameworks or automation for low-level software
  • Fluent in Python
  • Either proven Rust experience or a strong systems background (Go, C, C++) and a genuine appetite to build in Rust
  • Comfortable operating in a Kubernetes environment and reasoning about how software is built, containerized, deployed, and tested
  • Solid Linux systems background — the boot chain, kernel and drivers, low-level debugging — and comfort working close to the hardware
  • Real testing discipline: you've built automation that proves systems work, and you have strong opinions about flakiness, hermeticity, and signal
  • Clear communicator who treats the test framework as a product other engineers want to use, not a chore they route around
  • Must be eligible to access export controlled information (U.S. citizen, national, lawful permanent resident, refugee, or asylee, or eligible to access without authorization or obtain authorization)

Desired Qualifications

  • Rust and Kubernetes-native workflow orchestration (e.g., Argo Workflows) experience
  • HPC or large-cluster experience — InfiniBand/RoCE, GPU/accelerator validation, or performance-regression frameworks
  • Slurm or Slurm-on-Kubernetes (SUNK) experience
  • Firmware or lower-level hardware validation experience
  • Applying LLMs to test workflows — triage, failure classification, flaky-test detection
  • Contributions to open-source test frameworks, Rust crates, or systems projects

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce