Site Reliability Manager
On-siteChennai, Tamil Nadu, India
Job Summary
Oversee the day-to-day operations of the Site Reliability and Performance Engineering team, setting clear goals, supervising members, and providing technical leadership. Engage with cross-functional Product, Engineering, Security, Operations, Infrastructure teams, and vendors to improve mean time to detection and resolution. Design and implement infrastructure and application monitoring to ensure platform availability, while executing performance testing strategies including load, stress, and capacity planning using tools like JMeter, Locust, and LoadRunner. Automate performance testing within CI/CD pipelines to ensure continuous validation. Proactively analyze operational issues, innovate to improve efficiency, and achieve cost savings. Periodically review team performance and facilitate professional growth. As SRE and EIRE are global operational functions providing 24x7 support, weekend and public holiday coverage is an inherent expectation, offset through compensatory time off.
Required Qualifications
- Bachelor's degree in Computer Science or related field
- Minimum of 8 years of related experience
- Experience with Application Performance Management and Monitoring tools such as New Relic, AppDynamics, SiteSpect, and Datadog
- Experience with Content Delivery: Akamai
- Experience with Infrastructure monitoring tools like Zabbix, and Prometheus
- Experience with Databases eg: MongoDB, Oracle, Couchbase, Redis, MySQL
- Experience with Frameworks such as Dust/Angular, Nodejs, Springboot
- Experience with Log Analytics tools like Splunk, and ELK/Elastic
- Experience with Digital experience tools like Fullstory
- Experience with Performance Testing tools such as JMeter, Loadrunner, etc.
- Performance tuning experience with Tomcat, Node.js and Spring Boot
- Strong understanding of non-functional requirements, performance testing processes, and defect tracking
- 8+ years of experience with Cloud Technologies, at least half of which should be on the Microsoft Azure platform
- Strong hands-on experience with infrastructure and services (systems, network, cloud technology, provisioning, storage, etc)
- Must have strong experience with programming in one or more scripting languages (Python, Azure CLI, or Powershell)
- Hands-on experience with tool sets related to automation, orchestration, and managing infrastructure (Terraform, Puppet, Ansible, or Jenkins)
- Experience with configuring, deploying, and administering infrastructure and application monitoring tools that assist in troubleshooting performance and stability issues in a cloud environment
- Weekend and public holiday coverage
Desired Qualifications
- Master's degree in Computer Science Software Engineering or a related field
- Certifications in project management or specific software development methodologies
- Experience in working with cross-functional teams and stakeholders at high organizational levels
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.