Senior Systems Software Engineer - Fleet Debuggability
$184,000–$287,500 year
On-siteSanta Clara, California, United States
Job Summary
Architect, design, and build fleet-wide log collection and analysis solutions that aggregate signals across components, trays, and racks. Develop tooling to collect, normalize, and time-align logs from heterogeneous sources—including kernel, driver, syslog, Redfish, SEL, firmware, and BMC—over both in-band and out-of-band channels. Build and maintain a log catalog mapping raw signatures to fault classes and remediation guidance, then develop debug tooling that converts high-volume fleet logs into ranked, actionable diagnoses for hardware and platform faults. Drive design for low-overhead log analysis while partnering with developers, SWQA, and product engineering to define event schemas and manage open-source releases. Write design docs, own end-to-end delivery from definition through customer support, and perform code reviews to strengthen testing coverage.
Required Qualifications
- 10+ years in the software industry with specialization in system software and/or firmware development
- BS, MS, or PhD in CS, CE, EE, or a related technical field — or equivalent experience
- Proven track record of shipping scalable server products or fleet-wide experience
- A self-starter who loves finding creative solutions to complicated problems
- Excellent written and oral communication skills — including executive-level reporting
- Strong work ethic
- Dedication to teamwork
- Flexibility to work and communicate effectively across teams, partners, and time zones
- Experience with SCM (e.g., Git, Perforce) and project-management tools like Jira
- Strong, demonstrable skills in Python or RUST
- Deep Linux systems experience: kernel and driver logs, syslog, journald, and the realities of debugging on server platforms
- Hands-on experience with out-of-band management and platform interfaces — BMC, Redfish, IPMI, SEL — and an understanding of in-band vs. out-of-band trade-offs
- Strong skills in log parsing, normalization, and structured logging, and comfort designing schemas and taxonomies for machine-readable events
Desired Qualifications
- Experience leading debuggability solutions on sophisticated rack-scale compute architectures like GB200/GB300 NVL72
- Familiarity with log and telemetry analytics stacks (e.g., OpenSearch/ELK, Loki, Prometheus, Grafana, PagerDuty) and time-series databases
- Hands-on experience with x86/ARM system architecture and coding (C/C++, Python)
- Track record of integrating AI/LLM tooling into engineering workflows — for triage, validation, log analysis, or test generation
- Experience standing up follow-the-sun support organizations with measurable response SLAs
- Experience contributing to or maintaining open-source projects, including managing the boundary between internal and public code
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.