SRE Engineering Manager - GPU Cloud
HybridParis, Île-de-France, France
Job Summary
Lead and manage a team of 6 Site Reliability Engineers, supporting their career growth and technical execution within the GPU Cloud organization. Design and implement automated solutions for server lifecycle management across GPU clusters, alongside observability, logging, and monitoring solutions for large-scale infrastructure. Plan, prioritize, and manage the technical development roadmap while collaborating closely with software engineering, product, and cross-functional teams. Handle recruitment and career management for team members, and maintain, scale, and optimize high-availability production systems under heavy load. Participate in on-call rotations to ensure production reliability and fast incident resolution.
Required Qualifications
- Strong experience managing engineering teams in high-constraint production environments
- Proven expertise with Kubernetes container orchestration
- Direct experience with cluster management and virtualization tools (Proxmox, Warewulf)
- Experience with monitoring, metrics, and observability stacks (Prometheus, Grafana)
- Exposure to modern GPU hardware ecosystems (Nvidia, AMD) and high-speed networking fabric (InfiniBand, Spectrum-X, Tomahawk)
- Knowledge of distributed and high-performance storage solutions (Lustre DDN, VAST)
- Ability to handle high-pressure operational situations and manage incident stress pragmatically
- Excellent communication skills with the ability to convey challenging messages effectively
- Collaborative mindset with a focus on empowering engineers rather than micromanaging
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.