Senior Site Reliability Engineer
$160,000–$200,000 year
HybridSomerville, Massachusetts, United States
Job Summary
Mentor engineering teams on observability best practices, SLIs, and SLOs while contributing to triage and remediation processes. Perform incident response and debug production issues across the full stack, including TS/Go services on Kubernetes. Design and maintain core infrastructure and tooling for all engineering teams using Grafana, Loki, Mimir, Tempo, and Prometheus. Debug distributed systems using OpenTelemetry and manage metrics pipelines at scale with Prometheus and promQL. This role supports Tulip's mission to drive digital transformation in industrial environments through AI-native frontline operations solutions.
Required Qualifications
- 5+ years of experience working with open source Observability tools (e.g. Loki, Grafana, Tempo, Mimir stack)
- Hands-on experience instrumenting distributed systems using OpenTelemetry and managing metrics pipelines with Prometheus at scale
- Direct experience developing and distributing Claude Skills, Gemini Gems, or any other generic AI processes and are able to iterate on their efficacy
- Experience working with time-series data, ideally using promQL
- Must be able to lift 50 lbs
Desired Qualifications
- You can reason about systems at scale: their edge cases, failure modes, and life cycles across
- You're excited about setting the technical agenda and coming up with novel, broad ideas
- You regularly keep up with the newest AI advancements in the realm of Observability & Monitoring
- You know what a good SLA looks like, and can teach others how to spot one
- You can communicate as well as you can code. You understand the value of discussion and work best in a team that champions clear and frequent communication
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.