Truist logo
TruistPosted 2 weeks ago

Site Reliability Engineering Lead

On-siteAtlanta, Georgia, United States or Charlotte, North Carolina, United States

Full TimeSenior LevelEnterprise

Job Summary

Lead major incident responses and drive problem management to closure, establishing standardized playbooks and escalation paths. Architect automation solutions to eliminate toil and implement intelligent alerting leveraging AI and AIOps tools. Enhance telemetry coverage across logs, metrics, and traces using Dynatrace and Splunk, while defining enterprise observability practices and KPIs. Coach and mentor SRE team members to build technical depth, and partner with Delivery, Architecture, and Security teams to embed resilience into design. Lead large-scale reliability initiatives and workshops to elevate operational maturity across Wholesale. This senior technical leader drives improvements in automation, observability, and incident management within Truist's hybrid cloud and on-premises environments.

Required Qualifications

  • English (Required)
  • 1st shift (United States of America)
  • 7+ years of experience
  • expertise in distributed systems
  • Kubernetes
  • automation scripting
  • strong leadership in incident management
  • Bachelor's degree in Computer Science, Software Engineering, or related field
  • Minimum of 7 years of professional experience in software development
  • Deep knowledge of multiple programming languages, software architecture, and design principles
  • Deep understanding of software development lifecycle, testing, deployment, and security practices
  • Candidate must be willing to work onsite Monday - Friday at either office in Charlotte NC, Raleigh NC, or Atlanta, GA
  • Truist is a Drug Free Workplace
  • Truist is an Equal Opportunity Employer

Desired Qualifications

  • Advanced degree in Computer Science or related technical discipline
  • Professional certifications such as Certified Software Development Professional (CSDP) or equivalent
  • Deep expertise in cloud-native architectures, microservices, container orchestration, and DevOps
  • Strong familiarity with Agile frameworks, continuous integration/continuous deployment (CI/CD), and enterprise innovation management
  • 7+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or Infrastructure Operations
  • Deep hands-on experience with distributed systems, container orchestration (Kubernetes), and cloud-native operational tooling
  • Proficiency with automation and scripting languages (Python, Go, PowerShell, Ansible)
  • Strong understanding of observability platforms (Splunk, Dynatrace) and event-driven monitoring
  • Proven leadership in major incident management and cross-team technical coordination
  • Strong grasp of networking, Linux/Unix internals, and modern infrastructure patterns
  • Excellent communication skills, including executive-level situational awareness during critical incidents
  • Demonstrated ability to influence technical roadmaps and drive adoption of reliability best practices
  • Financial services or regulated industry experience
  • Experience enabling large-scale SRE transformations or modernization initiatives
  • Familiarity with chaos engineering, resilience assessments, and service failure modeling
  • Exposure to hybrid-cloud and multi-cloud operational frameworks
  • Experience contributing to or leading Center for Enablement functions or Communities of Practice

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce