Senior Observability Platform Engineer (80-100%)
$80,000–$100,000 year
HybridZürich, Zurich, Switzerland
Job Summary
Configure, operate, and enhance observability platforms including Clickhouse, Thanos, Loki, Tempo, and OpenTelemetry Collector to manage over 20 TB of daily telemetry from 10,000+ nodes. Drive organization-wide adoption of best practices by developing automated solutions for monitoring, alerting, and incident response while collaborating with engineering teams to optimize system performance and ensure high availability. Implement cost optimization strategies, capacity planning, and performance tuning using Kubernetes, Helm, and GitOps practices. Experiment with new tools and agentic AI workflows to accelerate platform capabilities and reduce alert noise.
Required Qualifications
- Proven track record in managing at least one of the following observability stacks: Thanos, Mimir, Cortex, Tempo, Loki or Clickhouse; with the ability to configure, operate, and improve these systems.
- Deep understanding of Kubernetes architecture and hands-on experience in managing resources on clusters.
- Experience in writing and maintaining Helm charts, and understanding third-party charts to deploy and manage Kubernetes resources efficiently.
- Experience in continuous delivery and GitOps practices (version control, CI/CD pipelines).
- Hands-on experience using agentic AI workflows (e.g., GitHub Copilot, Claude Code, Cursor, or similar) to accelerate day-to-day engineering.
- Expertise in containerization, orchestration, and optimization of Docker workloads.
- Bachelor's degree in Computer Science, Information Technology, or related field (or equivalent experience).
- 5+ years of experience in platform engineering, site reliability engineering, or a related role.
- Demonstrated experience in managing large-scale infrastructures and observability platforms (such as Thanos, Mimir, Cortex, Tempo, Loki, Clickhouse).
- You are excited by the prospect of managing more than 20 TB of telemetry data per day, originating from a fleet of 10 000+ nodes (including linux hosts, k8s clusters, VMs).
Desired Qualifications
- Coding knowledge in Golang or a similar language.
- Contributor to open source project written in Golang or a similar language.
- Knowledge of the OpenTelemetry Collector or direction contribution to project.
- Interest in applying AI/ML to the observability domain like anomaly detection on metrics and logs, automated root-cause analysis, alert noise reduction and correlation, and natural-language querying over telemetry
- Ability to quickly grasp new concepts and technologies, adapting to the evolving needs of the organization.
- Excellent communication skills, with the ability to convey complex technical concepts to both technical and non-technical stakeholders.
- Keen awareness of customer needs and the impact of platform operations on both internal engineering teams and external users.
- Strong ability to work collaboratively in cross-functional teams, contributing to a culture of continuous improvement and innovation.
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.