NVIDIA logo
NVIDIAPosted 1 month ago

Senior Systems Software Engineer - Fleet Debuggability

$184,000–$287,500 year

On-siteSanta Clara, California, United States

Full TimeSenior LevelEnterprise

Job Summary

Architect, design, and build fleet-wide log collection and analysis solutions that aggregate signals across components, trays, and racks. Develop tooling to collect, normalize, and time-align logs from heterogeneous sources—including kernel, driver, syslog, Redfish, SEL, firmware, and BMC—over both in-band and out-of-band channels. Build and maintain a log catalog mapping raw signatures to fault classes and remediation guidance, then develop debug tooling that converts high-volume fleet logs into ranked, actionable diagnoses for hardware and platform faults. Drive design for low-overhead log analysis while partnering with developers, SWQA, and product engineering to define event schemas and manage open-source releases. Write design docs, own end-to-end delivery from definition through customer support, and perform code reviews to strengthen testing coverage.

Required Qualifications

  • 10+ years in the software industry with specialization in system software and/or firmware development
  • BS, MS, or PhD in CS, CE, EE, or a related technical field — or equivalent experience
  • Proven track record of shipping scalable server products or fleet-wide experience
  • A self-starter who loves finding creative solutions to complicated problems
  • Excellent written and oral communication skills — including executive-level reporting
  • Strong work ethic
  • Dedication to teamwork
  • Flexibility to work and communicate effectively across teams, partners, and time zones
  • Experience with SCM (e.g., Git, Perforce) and project-management tools like Jira
  • Strong, demonstrable skills in Python or RUST
  • Deep Linux systems experience: kernel and driver logs, syslog, journald, and the realities of debugging on server platforms
  • Hands-on experience with out-of-band management and platform interfaces — BMC, Redfish, IPMI, SEL — and an understanding of in-band vs. out-of-band trade-offs
  • Strong skills in log parsing, normalization, and structured logging, and comfort designing schemas and taxonomies for machine-readable events

Desired Qualifications

  • Experience leading debuggability solutions on sophisticated rack-scale compute architectures like GB200/GB300 NVL72
  • Familiarity with log and telemetry analytics stacks (e.g., OpenSearch/ELK, Loki, Prometheus, Grafana, PagerDuty) and time-series databases
  • Hands-on experience with x86/ARM system architecture and coding (C/C++, Python)
  • Track record of integrating AI/LLM tooling into engineering workflows — for triage, validation, log analysis, or test generation
  • Experience standing up follow-the-sun support organizations with measurable response SLAs
  • Experience contributing to or maintaining open-source projects, including managing the boundary between internal and public code

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce