Principal Resiliency Engineer
HybridChennai, Tamil Nadu, India
Job Summary
Build and deploy prototype applications that showcase observability capabilities across hybrid environments. Conduct rigorous testing with simulated disruptions, environmental failures, and performance scenarios to validate proposed standards. Collaborate with platform teams, application owners, and external vendors to integrate observability into real workloads. Produce runbooks, configuration guides, architectural patterns, reusable dashboards, and findings to support enterprise enablement. Provide technical feedback to influence enterprise strategy and host sprint retrospectives for stakeholders. Evaluate proposed standards for distributed tracing, metrics, logging, and alerting. Join the Reliability Architecture team to shape enterprise observability and support platform modernization.
Required Qualifications
- Minimum of 8 years in distributed application design and implementation
- Bachelor's degree in computer engineering or equivalent experience
- 5+ years in enterprise Java technologies and open standards
- 5+ years in infrastructure, networking, middleware, and database architecture
- 5+ years in highly available architecture and disaster recovery
- 3+ years in containers and cloud-based solution delivery
- Strong troubleshooting and performance analysis skills
- Java
- Python
- Bash
- SQL
- CI/CD pipelines and automation frameworks (e.g., Jenkins, Selenium)
- OpenTelemetry, distributed tracing, metrics generation, logging
- Chaos engineering tools (e.g., Gremlin, AWS FIS)
- AWS and Azure cloud environments
- Infrastructure as Code (IaC), container orchestration, hybrid deployments
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.