Encora logo
EncoraPosted 4 weeks ago

Chaos Test Lead

On-siteMexico City, Mexico City, Mexico

Full TimeSenior LevelMediumTechnology

Job Summary

Lead and execute the chaos engineering lifecycle, including planning, designing, and reporting experiments to ensure system resilience. Schedule, staff, and document recovery and resilience testing while managing remediation and closure of identified issues. Analyze architecture to recommend weak areas prone to failure and collaborate with business and enterprise teams to define high availability requirements. Design and automate continuous chaos experiments using tools like Gremlin, Chaos Native, or Litmus within Unix/Linux and cloud environments (AWS, GCP, Azure). Troubleshoot failures in CI/CD pipelines and diagnose complex distributed systems.

Required Qualifications

  • 5+ years
  • Relevant experience on Chaos engineering / Resilience / High availability testing
  • atleast 5-10 years of relevant experience
  • Chaos Test Planning
  • Chaos Test Designing
  • Reporting
  • Implement and lead execution of the chaos engineering Lifecycle
  • Ensure recovery and resilience testing is scheduled, staffed, executed, and documented, including remediation and closure of issues
  • Ability to analyse the architecture & recommend weak areas that are likely to failure / outages
  • Ability to work with Business & technology teams to identify and report on resilience / High availability requirements
  • Ability to work with enterprise architecture and development teams to architect applications for high availability and resiliency
  • Design, develop and execute automated / continuous Chaos Engineering experiments
  • Ability to troubleshoot the failures in CI/CD pipeline
  • Automate Chaos experiments through chaos engineering tools (Gremlin / Chaos Native / Litmus etc) to run continuously
  • Hands on experience in Unix/Linux OS environments and operating system internals, file systems, disk/storage and networking protocols
  • Strong knowledge on Public cloud platforms – AWS, GCP, Azure
  • Knowledge on Monitoring, Alerting, Logging
  • Knowledge on VPC's, proxy's, load balancers, availability zones
  • Ability in diagnosing and debugging complex distributed systems
  • Tools (any of these) - Gremlin, Chaos Native, Litmus
  • Strong leadership skills and ability to work in a cross-functional environment
  • Strong interpersonal, oral, and written communication skills
  • Strong analytical and decision-making skills

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce