Lead Site Reliability Engineer (Performance & Scalability) | Contract | Remote MX
RemoteMexico
Job Summary
Establish performance baselines, define SLOs and error budgets, and instrument the full request path across application services, databases, and third-party dependencies. Lead cross-functional remediation for identified bottlenecks, build capacity models, and execute load, stress, soak, and failure testing in representative environments. Drive architecture hardening, resilience improvements, and automated performance gates while creating operational runbooks for scale-up events and incidents. Translate technical risks into business implications for leadership and recommend capacity investments before constraints emerge. Own technical readiness assessments for major pilots and production launches.
Required Qualifications
- Significant experience in Site Reliability Engineering, performance engineering, platform engineering, distributed systems, or a closely related engineering discipline
- Experience supporting production systems with meaningful scale, traffic, latency, or availability requirements
- Deep understanding of observability, performance analysis, capacity planning, and reliability engineering
- Strong hands-on experience with cloud infrastructure and production distributed systems
- Deep knowledge of databases, networking, caching, queueing, compute, storage, and common distributed-system failure modes
- Experience defining and operating against SLOs, SLIs, error budgets, and production reliability metrics
- Hands-on experience performing load, stress, soak, scalability, and resilience testing
- Ability to profile systems, diagnose bottlenecks, tune architecture, and work directly with engineering teams to implement improvements
- Experience designing for graceful degradation, dependency failures, recovery, and high-demand scenarios
- Strong incident management and root-cause analysis experience
- Ability to translate technical performance and reliability risks into clear business implications for senior leadership
- Strong judgment around when systems genuinely require optimization versus when additional complexity is premature
Desired Qualifications
- Experience operating high-scale SaaS, identity, DNS, registry, infrastructure, or other highly distributed platforms
- Experience creating capacity-cost models and forecasting infrastructure requirements
- Experience building performance and reliability gates into CI/CD pipelines
- Experience preparing platforms for significant increases in traffic associated with enterprise customers or strategic partnerships
- Experience leading reliability or performance initiatives that span multiple engineering teams
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.