Scaleway logo
ScalewayPosted 1 month ago

SRE Engineering Manager - GPU Cloud

HybridParis, Île-de-France, France

Full TimeStartup

Job Summary

Lead and manage a team of 6 Site Reliability Engineers, supporting their career growth and technical execution within the GPU Cloud organization. Design and implement automated solutions for server lifecycle management across GPU clusters, alongside observability, logging, and monitoring solutions for large-scale infrastructure. Plan, prioritize, and manage the technical development roadmap while collaborating closely with software engineering, product, and cross-functional teams. Handle recruitment and career management for team members, and maintain, scale, and optimize high-availability production systems under heavy load. Participate in on-call rotations to ensure production reliability and fast incident resolution.

Required Qualifications

  • Strong experience managing engineering teams in high-constraint production environments
  • Proven expertise with Kubernetes container orchestration
  • Direct experience with cluster management and virtualization tools (Proxmox, Warewulf)
  • Experience with monitoring, metrics, and observability stacks (Prometheus, Grafana)
  • Exposure to modern GPU hardware ecosystems (Nvidia, AMD) and high-speed networking fabric (InfiniBand, Spectrum-X, Tomahawk)
  • Knowledge of distributed and high-performance storage solutions (Lustre DDN, VAST)
  • Ability to handle high-pressure operational situations and manage incident stress pragmatically
  • Excellent communication skills with the ability to convey challenging messages effectively
  • Collaborative mindset with a focus on empowering engineers rather than micromanaging

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce