Data Center Operations and Maintenance Engineer
HybridBellevue, Washington, United States
Job Summary
Monitor live data center and GPU infrastructure for performance, capacity, and reliability issues. Lead or support incident response for production issues, driving toward fast, effective resolution. Collaborate with engineering teams to ensure monitoring, alerting, and operational tooling are first class so issues are caught before impacting customers. Partner with hardware, networking, orchestration, and infrastructure teams to resolve root causes and prevent recurrence. Contribute to runbooks, on-call practices, and operational maturity as the platform scales. This individual contributor role helps build and scale the Operations & Maintenance function supporting mission-critical AI infrastructure. You will partner closely with Engineering, Infrastructure, Networking, Hardware, and Operations and Maintenance leadership to create a highly reliable, scalable operating system capable of supporting one of the industry's most advanced AI infrastructure platforms.
Required Qualifications
- Experience in a data center operations, site reliability, or infrastructure operations role
- Experience supporting GPU or large-scale compute environments
- Strong incident response and troubleshooting skills across hardware, networking, and systems layers
- Comfortable working in an early-stage environment where processes and tooling are still being established
- U.S. work authorization
- Hybrid role based in the Bellevue, WA area
- Approximately three days per week in the office
Desired Qualifications
- Experience in multiple technical environments
- Experience with on-call/incident-management tooling (e.g., PagerDuty, Opsgenie) and monitoring stacks (e.g., Prometheus, Grafana, Datadog)
- Background supporting GPU cluster or data center operations post-launch
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.