AT&T logo
AT&TPosted 1 week ago

Principal System Engineering - SRE

$155,400–$261,100 year

On-siteDallas, Texas, United States or Plano, Texas, United States

Full TimeSenior LevelEnterprise

Job Summary

Analyze production incidents end-to-end across applications, infrastructure, and cloud environments using observability data to identify root causes, patterns, and systemic weaknesses. Turn incident insights into high-quality postmortems and partner with engineering teams to drive corrective actions and long-term improvements. Shift the organization from reactive response to proactive reliability by combining system-level thinking with data, automation, and AI-assisted analysis to implement permanent fixes and preventive measures. This role focuses on eliminating recurring issues rather than just fixing incidents, requiring deep RCA skills and strong communication to explain complex problems. You will work within the Corporate Systems Reliability and Software Delivery teams to advance information technology performance and maximize ROI.

Required Qualifications

  • Office presence of a minimum of 5 days per week
  • Location in the location(s) posted
  • 7+ years in Systems Engineering, ITSM, RM/CM
  • Background in SRE, Support or QA
  • One or more of the following SRE Tools: T-APM, T-Trace, CatchPoint, Grafana
  • Hands-on experience and understanding of concepts and tools such as SAFe, Agile, DevOps, CI/CD, Data Analytics, and building Gen AI use cases
  • Experience with AI technologies, Python, SQL, data analytics, Power BI and ITSM tools (e.g., ServiceNow)
  • Modern Enterprise Release Management/Change Management and ITSM
  • BS/BA in Computer Science
  • Preferred tools: modern Release Management processes for Agile and DevOps environments
  • Jira Align, JSM, Jira Cloud, Git for enterprise RM/CM
  • Relevant certifications (SAFe, Agile, DevOps, AI/ML)

Desired Qualifications

  • Background in QA, test engineering, or automation engineering (strong plus)
  • Experience using AI or advanced analytics for incident analysis or pattern detection
  • Understanding of distributed systems and failure modes
  • Experience with data analysis / visualization tools (e.g., Power BI, Tableau)

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce