Staff Data Center Operations Engineer
$150,000–$170,000 year
On-siteDenver, Colorado, United States
Job Summary
Own Tier 2/3 hardware escalations across all Crusoe sites for issues exceeding local capability, engaging directly with OEM and ODM engineering teams to drive resolution. Travel to sites for complex platform issues, new hardware bring-ups, and deployment support while identifying recurring failure patterns to translate them into platform feedback, sparing strategy inputs, or OEM improvement requests. Root-cause complex hardware issues including PCIe, BMC, and thermal faults, then hand off findings to internal engineering teams with clear, well-documented escalation packages. Develop and maintain deep technical relationships with SuperMicro and HPE partners, serving as the technical voice in conversations to influence hardware roadmaps and drive fleet-wide improvements. Own the development of platform-specific SOPs, runbooks, and field troubleshooting procedures, and deliver technical training for SiteOps technicians covering hardware architecture and deployment procedures. Serve as SiteOps' senior technical representative at Denver headquarters, partnering with engineering and procurement on sparing strategy and RMA lifecycle management while providing operational input into next-generation GPU platform evaluations.
Required Qualifications
- 7+ years in data center operations, field engineering, or OEM/ODM technical support with hands-on GPU infrastructure experience
- Direct hands-on experience deploying and supporting GPU platforms at scale across one or more major OEMs or ODMs
- Familiarity with SuperMicro and HPE platforms required
- Deep familiarity with server platform architecture and OEM escalation and RMA processes
- Experience leading or contributing to large-scale GPU cluster bring-ups including rack staging and production handoff
- Demonstrated ability to build technical relationships with OEM and ODM engineering teams and drive platform-level issue resolution
- Experience developing SOPs, runbooks, or field troubleshooting procedures and delivering technical training to data center technician teams
- Strong written communication — comfortable producing escalation documentation, platform runbooks, and leadership reporting
- Willingness to travel domestically and internationally to Crusoe sites as needed (target: up to 30%)
Desired Qualifications
- Direct experience with SuperMicro GPU platforms (B200, GB200, or newer)
- SuperMicro Certified Engineer credentials
- Familiarity with ASUS or Quanta server platforms and ODM engagement models
- Experience with liquid-cooled GPU platforms and CDU integration
- Familiarity with AMD Instinct GPU platforms (MI300X/MI350X/MI355X)
- Prior experience at an AI cloud provider, hyperscaler, or GPU-first infrastructure operator
- Experience contributing to technician certification programs or IC leveling standards within a DC ops organization
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.