Designworks Talent logo
Designworks TalentPosted 1 month ago

Data Center Operations and Maintenance Engineering Leader

$31,200–$31,200 year

HybridBellevue, Washington, United States

Full TimeSenior LevelStartup

Job Summary

Build and lead high-performing Operations & Maintenance teams while developing the operational strategy and organizational structure for large-scale AI infrastructure. Establish world-class processes for incident response, change management, and service reliability; lead major incident management and executive communications during production events. Drive operational excellence through proactive monitoring, observability, automation, and continuous improvement initiatives. Partner with Engineering to ensure readiness for new deployments, define SLOs and reliability metrics, and build scalable on-call programs and operational governance. Champion root cause analysis to improve platform resilience and influence infrastructure architecture for better availability and efficiency.

Required Qualifications

  • Experience leading Operations, Site Reliability, Infrastructure Operations, Data Center Operations, or Production Engineering organizations
  • Proven success building or scaling operations teams within cloud infrastructure, hyperscale environments, AI infrastructure, or large distributed systems
  • Deep expertise in production operations, incident management, service reliability, and operational excellence
  • Experience leading cross-functional teams during high-severity production incidents
  • Strong understanding of infrastructure operations across compute, networking, storage, and hardware environments
  • Demonstrated success building operational processes, organizational structure, and scalable support models in high-growth environments
  • Executive-level communication skills with the ability to influence engineering and business leadership
  • Comfortable operating in an early-stage organization where many systems and processes are being built for the first time
  • U.S. work authorization
  • Hybrid role based in the Bellevue, WA area
  • Approximately three days per week in the office

Desired Qualifications

  • Experience supporting hyperscale cloud platforms, GPU infrastructure, AI platforms, HPC environments, or large-scale data centers
  • Experience with modern observability and monitoring platforms such as Grafana, Prometheus, Datadog, or similar technologies
  • Familiarity with incident management platforms including PagerDuty, Opsgenie, or equivalent solutions
  • Experience implementing operational maturity frameworks and reliability engineering best practices
  • Track record of building globally distributed operations organizations

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce