Lead Site Reliability Engineer - Corporate KYC
On-siteHouston, Texas, United States
Job Summary
Conduct resiliency design reviews and act as the main point of contact during major incidents to identify and solve issues quickly. Lead initiatives to improve application reliability using data-driven analytics, establish service level objectives with stakeholders, and champion site reliability culture across the team. Guide medium to large-sized products by breaking complex problems into digestible work, providing technical mentorship, and utilizing enterprise-authorized AI for incident triage and workflow automation. Document knowledge internally and apply SRE best practices in domains like observability, container orchestration, and CI/CD.
Required Qualifications
- Formal training or certification in software engineering concepts with 5+ years of applied experience
- Deep proficiency in reliability, scalability, performance, security, enterprise system architecture, toil reduction, and other site reliability best practices with the ability to implement these practices within an application or platform
- Fluency in at least one programming language such as (e.g., Python, Java Spring Boot, .Net, etc.)
- Demonstrable experience using enterprise-authorized AI capabilities within the work environment to improve SRE workflows (e.g., incident investigation support and knowledge capture) with strong validation habits and awareness of data sensitivity
- Ability to evaluate AI-assisted operational recommendations for correctness and risk, define appropriate guardrails for team usage, and ensure outcomes align to resiliency and security expectations
- Deep knowledge of software applications and technical processes with emerging depth in one or more technical disciplines
- Proficiency and experience in observability such as white and black box monitoring, SLO alerting, and telemetry collection using tools such as Grafana, Dynatrace, Prometheus, Datadog, Splunk, etc.
- Proficiency in continuous integration and continuous delivery tools (e.g., Jenkins, GitLab, Terraform, etc.)
- Experience with container and container orchestration (e.g., ECS, Kubernetes, Docker, etc.)
- Experience with troubleshooting common technologies and issues
Desired Qualifications
- Ability to teach new programming languages to team members
- Ability to expand and collaborate across different levels and stakeholder groups
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.