Senior Site Reliability Engineer, DGX Cloud
RemoteZürich, Zurich, Switzerland or Switzerland
Job Summary
Build, implement, and support operational and reliability aspects of large-scale Kubernetes clusters with a focus on performance at scale, real-time monitoring, logging, and alerting. Define SLOs/SLIs, monitor error budgets, and streamline reporting while maintaining services through system creation consulting, capacity management, and launch reviews. Operate and optimize GPU workloads across AWS, GCP, Azure, OCI, and private clouds, scaling systems sustainably via automation. Lead triage and root-cause analysis of high-severity incidents, practicing balanced incident response and blameless postmortems. Participate in on-call rotation to support production services.
Required Qualifications
- BS in Computer Science or related technical field, or equivalent experience
- 10+ years of experience operating production services
- Expert-level knowledge of Kubernetes administration, containerization, and microservices architecture
- Experience with infrastructure automation tools (e.g., Terraform, Ansible, Chef, Puppet)
- Proficiency in at least one high-level programming language (e.g., Python, Go)
- In-depth knowledge of Linux operating systems, networking fundamentals (TCP/IP), and cloud security standards
- Proficient knowledge of SRE principles, encompassing SLOs, SLIs, error budgets, and incident handling
- Experience building and operating comprehensive observability stacks (monitoring, logging, tracing) using tools like OpenTelemetry, Prometheus, Grafana, ELK Stack, Lightstep, Splunk, etc.
Desired Qualifications
- Operating GPU-accelerated clusters with KubeVirt in production
- Applying generative-AI techniques to reduce operational toil
- Experience with workflow orchestration platforms such as Temporal, Cadence, Airflow, Argo Workflows, or Step Functions
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.