Staff Software Engineer, Quality & Reliability
$172,000–$229,000 year
HybridToronto, Ontario, Canada
Job Summary
Define and drive BuildOps' technical strategy for engineering quality, production reliability, and safe software delivery. Identify systemic sources of customer-impacting failures and lead cross-team initiatives that address root causes rather than individual symptoms. Establish architectural principles, engineering standards, and paved roads that make reliable system design and safe delivery easier by default. Partner with engineering teams during system and product design to improve resilience, operability, testability, and failure isolation before implementation begins. Build or guide the development of shared platform capabilities for release safety, automated validation, production feedback, test data, environment management, and developer self-service. Advance BuildOps' observability strategy so teams can understand system behavior, detect regressions quickly, diagnose failures, and make informed reliability investments. Mentor engineers and technical leaders, raising the organization's ability to reason about risk, reliability, and quality throughout the software lifecycle. Evaluate BuildOps' existing practices and technology objectively, evolving or replacing them when they no longer meet our needs.
Required Qualifications
- Significant software engineering experience
- Operating at Staff or equivalent scope on ambiguous, cross-cutting technical problems
- A track record of leading multi-team initiatives that improved production reliability, software delivery, platform capabilities, or engineering effectiveness
- Strong systems thinking
- Experience designing and operating distributed systems in a cloud environment such as AWS
- Strong software design and programming skills in TypeScript, Java, or another relevant language
- Experience with several of the following: observability, resilience engineering, CI/CD, release safety, automated validation, developer platforms, testability, performance engineering, or incident learning
- The ability to define useful engineering measures while avoiding metrics that reward activity without improving outcomes
- Demonstrated success influencing architecture and engineering practices across teams that do not report to you
- Strong written and verbal communication, including the ability to explain technical risks, tradeoffs, and strategy to engineering, product, and business stakeholders
- A practical approach that balances long-term direction with incremental improvements that deliver value quickly
- Comfortable moving between long-term technical strategy and hands-on implementation when necessary
- Build credibility through sound judgment, clear communication, and measurable results
- Care deeply about the experience of both our customers and the engineers building for them
- Must be able to support employment in the U.S. where BuildOps is registered to do business
- Not located in Alaska, Hawaii, Kentucky, Mississippi, Nebraska, New Mexico, North Dakota, Rhode Island, South Dakota, West Virginia, or Wyoming
Desired Qualifications
- Experience with several of the following: observability, resilience engineering, CI/CD, release safety, automated validation, developer platforms, testability, performance engineering, or incident learning
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.