Software Engineer, Observability
$108,000–$135,000 year
HybridToronto, Ontario, Canada
Job Summary
Maintain, improve, and develop tooling that enhances platform reliability, scalability, and efficiency. Assist engineering teams in defining service-level objectives (SLOs) and monitor feature development speed against reliability. Analyze metrics from operating systems, control planes, and applications to support fault detection and performance enhancement. Collaborate with cross-functional teams to align observability with design reviews and capacity planning. Automate repetitive tasks and keep documentation at a world-class level. Participate in on-call rotations to respond to incidents and mitigate customer-impacting events. This role supports Lyft's purpose of serving and connecting users through robust infrastructure.
Required Qualifications
- 3+ years of experience working on teams responsible for software development, automation, and systems engineering
- Bachelor's Degree or equivalent experience in Computer Science or a relevant discipline
- Proficiency in creating production-ready code in one or more high-level languages, such as Go or Python
- Experience operating infrastructure in public cloud environments, such as AWS, including familiarity with Managed Services
- Experience in building and maintaining observability infrastructure to support robust monitoring and analysis
- Familiarity with Kubernetes and managing multi-cluster environments in production settings
- Proven track record with modern Observability stack, including proficiency in Prometheus, Grafana, Loki, and other open-source tracing and alerting frameworks
- Team Members will be expected to work in the office at least 3 days per week, including on Mondays, Wednesdays, and Thursdays
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.