Sr. Site Reliability Engineer
HybridPhiladelphia, Pennsylvania, United States
Job Summary
Senior Site Reliability Engineer role focused on ensuring highest availability and resiliency of a rapidly growing global payment platform. The role emphasizes AI-enabled operations, automation, and self-healing capabilities, with responsibilities spanning creating and improving observability, automating remediation, leading incident response, and reducing toil through AI-assisted tooling. The position participates in an on-call rotation, offers a hybrid schedule with potential remote arrangements for exceptional candidates, and requires collaboration across cross-functional teams. Candidates should bring 5+ years of relevant experience, strong problem-solving and communication skills, expertise in modern APM/AIOps tools (Dynatrace, Datadog, New Relic), proficiency with scripting (PowerShell/Python), SQL, networking fundamentals, container orchestration (Kubernetes/AKS), cloud and virtualization (Azure, VMware), and experience applying AI-driven automation to reliability and operations, while adhering to PCI and security obligations.
Required Qualifications
- BS degree in Computer Science or equivalent, or equivalent years of relevant experience
- Minimum of 5 years of hands-on technical experience in highly available, high-throughput, web-based technology environments
- Demonstrated history of self-directed learning
- Next-level problem-solving abilities and a strong bias toward practical, proven solutions
- A track record of identifying and eliminating manual toil through automation
- Excellent communication and organizational skills, with a strong sense of ownership and service
- Expert-level proficiency in an enterprise APM platform and its AI/ML-driven (AIOps) capabilities (Dynatrace strongly preferred)
- Hands-on experience with AI-assisted development and automation tools (Anthropic Claude, OpenAI Codex, Azure AI services) and applying them to operational work
- Proficiency in scripting and automation (PowerShell and/or Python)
- Strong SQL / T-SQL skills
- Solid understanding of core networking concepts (DNS, HTTP/HTTPS, load balancing, TCP/IP)
- Working knowledge of container orchestration, IaaS/PaaS cloud services (Azure), and VMware
- Working knowledge of application development processes
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.