Principal Platform Engineer (High Availability & Disaster Recovery)
On-siteGurugram, Haryana, India
Job Summary
Own the architectural strategy and technical execution for the platform's resilience, high availability, and disaster recovery lifecycle. Design, build, and deliver solutions supporting global availability, resiliency, and multi-region disaster preparedness between cloud and on-premises environments. Lead cross-functional coordination to define guardrails for system fault-tolerance, implement automated failover mechanisms, and champion chaos engineering practices to eliminate single points of failure. Establish the foundational Resiliency Engineering framework, set platform-wide standards for self-healing, and introduce the organization's first automated fault-injection and live-fire failover drills. Partner with engineering squads to embed recovery protocols and monitor infrastructure KPIs to mitigate risks before causing downtime while ensuring compliance with strict Customer SLA requirements.
Required Qualifications
- 8+ Years of HA Engineering
- Technical Influence
- Failover & Replication
- Infrastructure as Code (IaC)
- Systems & Container Engineering
- Hybrid Network Engineering
- Strategic Execution
Desired Qualifications
- Experience supporting SaaS products
- Experience with Incident Management, Post Mortems and related practices
- Knowledge of observability and monitoring best practices
- Experience operating within one or more public clouds (AWS, GCP, Azure)
- Experience with configuration management, and infrastructure as code
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.