Senior Manager, Site Reliability Engineering – Paylo Platform
On-siteHouston, Texas, United States or Dallas, Texas, United States
Job Summary
Directly manage and develop 3 SRE Managers/Leads while owning the health, growth, and performance of a ~20-person SRE organization supporting the Paylo product suite. Set the vision, priorities, and operating cadence for the SRE function, translating business and product goals into a reliability roadmap. Build a strong bench by hiring, coaching, and developing managers and senior engineers with clear career paths and succession plans. Foster a blameless, learning-oriented culture around incidents, on-call rotations, and operational excellence. Partner closely with engineering directors, product managers, and business stakeholders to align reliability investments with business risk and customer impact. Stay technically engaged by participating in architecture reviews, troubleshooting complex production issues, and contributing to infrastructure-as-code, Kubernetes manifests, and CI/CD pipelines. Set and enforce engineering standards for multi-cloud infrastructure across AWS and Azure, container orchestration on Kubernetes, and GitOps-based continuous delivery using Argo CD. Own the Infrastructure-as-Code strategy, CI/CD pipeline architecture, and observability strategy using Datadog as the standard platform. Define SLIs/SLOs, error budgets, and reliability KPIs while owning the end-to-end incident management program, including escalation paths, postmortems, and root-cause analysis. Ensure resilience, disaster recovery, and capacity planning practices for payment and transaction systems, while partnering with Security and Compliance to maintain PCI DSS readiness. Track and report cost, capacity, and operational KPIs to senior leadership.
Required Qualifications
- 8+ years of experience in Site Reliability Engineering, DevOps, or Infrastructure/Platform Engineering, including 4+ years in a people-leadership role
- Proven experience managing managers — you have directly led team leads/managers, not just individual contributors, and are comfortable operating at the scale of ~20 total reports
- Strong, hands-on expertise across AWS and Azure — you can architect, troubleshoot, and operate multi-cloud infrastructure yourself, not just direct others to do so
- Strong, hands-on expertise with Kubernetes and Helm — cluster operations, troubleshooting at scale, and chart design/maintenance
- Strong, hands-on expertise with Argo CD/Argo Workflows for GitOps-based continuous delivery
- Strong, hands-on expertise with Infrastructure as Code (Terraform, OpenTofu), including module design and state management
- Strong, hands-on expertise with Jenkins for CI/CD pipeline design, administration, and automation
- Strong, hands-on expertise with Datadog (or equivalent enterprise observability platform), including designing monitoring/alerting strategy, dashboards, and APM/tracing at scale
- Demonstrated track record of driving incident management, on-call, and postmortem programs for high-traffic, customer-facing systems
- Excellent communication and stakeholder-management skills; able to represent SRE to engineering leadership and business partners with equal credibility
- A strong, visible leadership style — someone who sets clear direction, holds teams accountable, and builds trust across the organization
- Applicants must be legally authorized to work in the United States without the need for employer sponsorship, now or in the future. PDI Technologies is unable to offer visa sponsorship for this role
Desired Qualifications
- Experience supporting payments, fuel/retail, or loyalty platforms, or other systems with PCI DSS or similar compliance obligations
- Relevant certifications such as CKA/CKAD, AWS Certified Solutions Architect, Microsoft Certified: Azure Solutions Architect, or HashiCorp Terraform Associate
- Experience with messaging systems (Kafka/SQS/SNS), PagerDuty (or similar), and multi-region/multi-AZ resilience patterns
- Prior experience consolidating or standardizing SRE and DevOps practices across multiple product lines or recently-integrated/acquired teams
- Experience partnering with product and business stakeholders to translate reliability investments into business outcomes
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.