Grafana Labs logo
Grafana LabsPosted 1 month ago

Staff Software Engineer - Databases SRE | Sweden | Remote

$878,578–$1,054,294 year

RemoteUnited Kingdom or Spain

Full TimeSenior LevelMediumTechnology

Job Summary

Own production reliability for high-SLA Grafana Cloud database environments based on Mimir, Loki, Tempo, and Pyroscope. Design and implement automation to scale reliability practices while defining per-tenant SLOs and proactively reducing SLO burn. Serve as the primary escalation point for incidents, leading customer-impacting response and post-incident reviews. Partner with product engineering squads to influence feature design for production scalability and operability. Improve observability within customer environments and develop fault-tolerant design patterns across the service lifecycle. Collaborate with engineering leaders to define product strategy and teach best practices to the team. This Staff Software Engineer role is embedded within the Mimir, Loki, and Tempo squads, supporting a 100% remote organization of 1,600+ members. Candidates must have 8+ years of engineering experience with 4+ in SRE/production environments.

Required Qualifications

  • 8+ years engineering experience
  • 4+ years in SRE/CRE/production engineering
  • Strong preference for those with formal customer reliability engineering experience
  • Strong Kubernetes experience in AWS, GCP, or Azure
  • Familiarity with infrastructure-as-code tooling (Helm, Terraform, Jsonnet, etc.)
  • Strong experience with technical leadership, leading a team through projects, mentoring other engineers on the team and serving as a force-multiplier
  • Experience operating multi-tenant systems in production
  • Strong experience designing and implementing SLOs
  • Experience with one or more programming languages (e.g. Go, Python, Java, etc)
  • Experience with Linux operating systems internals
  • Some knowledge of networking, cloud storage, and scaling
  • Excellent problem-solving and troubleshooting skills
  • Experience with calmly and actively participating in blame-free Incident Response, following up on actions, and writing high quality PIRs (Post Incident Reviews, a.k.a. post-mortem documents)
  • Ability to reason about performance, scaling, and failure modes
  • Comfortable working within an engineering team where individuals are encouraged to have a strong sense of autonomy and self-direction
  • Ability to partner deeply with product engineering teams
  • Candidates from the UK, Sweden, Spain or Germany

Desired Qualifications

  • Intellectual curiosity
  • Defaulting to transparency
  • High bias towards action
  • Kindness

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce