Hydra Host logo
Hydra HostPosted 1 month ago
EXPIRED

Operations and Support Lead

$140,000–$200,000 year

RemoteUnited States

Full TimeSenior Level

Job Summary

Design and lead Hydra Host's customer support organization from the ground up, interfacing with ML engineering, GPU infrastructure customers, and data center partners to deliver enterprise-grade service levels. Coordinate daily operational health across engineering, vendors, and customers to resolve GPU cluster failures, network outages, and hardware issues as Incident Commander. Build scalable support operations including ticketing systems, escalation procedures, and AI-powered tools while managing vendor relationships and enforcing SLAs. Establish operational KPIs, monitor MTTR and uptime, and provide executive reporting on performance. Act as the operational bridge between technical teams and customer-facing organizations to ensure the NeoCloud platform runs efficiently.

Required Qualifications

  • 5+ years leading customer support, technical operations, infrastructure operations, or service delivery organizations
  • Experience supporting enterprise infrastructure, cloud platforms, AI infrastructure, NeoCloud providers, or large-scale data center environments
  • Experience managing production incidents in mission-critical environments
  • Strong understanding of servers, networking, storage, and enterprise infrastructure
  • Experience building operational processes that scale rapidly
  • Experience working with third-party vendors and infrastructure providers
  • Excellent communication skills with both technical and executive stakeholders
  • Proven ability to lead cross-functional initiatives across engineering, operations, and customer organizations
  • Strong organizational and project management skills
  • Bias toward ownership, accountability, and continuous improvement

Desired Qualifications

  • Experience operating NeoCloud or AI Factory infrastructure
  • Experience supporting NVIDIA GPU environments (HGX, DGX, H100, H200, Blackwell)
  • Experience with bare-metal cloud infrastructure
  • Familiarity with InfiniBand, RoCE, high-performance networking, and AI storage platforms
  • Experience with ITIL Service Management, Incident Management, Change Management, and Problem Management
  • Experience implementing AI-driven support automation and operational tooling
  • Experience with CRM and service management platforms including Zendesk, Jira Service Management, Linear, HubSpot, Intercom, or similar platforms
  • Previous experience managing technical support or operations teams

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Find similar roles