DevOps Engineer
On-siteAtlanta, Georgia, United States
Job Summary
Proactively identify and remediate platform risks across critical payment processing services by working directly from post-incident findings and root cause analyses. Onboard services onto observability platforms like New Relic to establish dashboards and alerts, while owning Helm chart configuration for containerized applications on AWS Elastic Kubernetes Service. Build and maintain internal engineering tooling covering deployment pipelines, certificate lifecycle automation, and AI-powered developer tools to reduce toil. Participate in 24/7 on-call rotations to provide incident response support and partner with engineers during major incidents to drive durable remediation. This role is based in the Atlanta, GA office with a minimum three-day in-office requirement.
Required Qualifications
- 3+ years of hands-on DevOps or platform engineering experience in a production environment
- Strong proficiency with Kubernetes and Helm
- experience tuning workloads on managed Kubernetes services such as AWS Elastic Kubernetes Service
- Practical observability experience with New Relic or equivalent platforms such as Datadog or Prometheus
- able to build meaningful dashboards and alert configurations from scratch
- Proficiency in at least one scripting or programming language (Python, Go, TypeScript, or similar)
- a track record of building automation that other engineers actually use
- Familiarity with continuous integration and delivery pipelines
- GitOps practices
- infrastructure-as-code tooling such as Terraform or Ansible
- Experience managing TLS certificates
- a working understanding of public key infrastructure in production environments
- Comfort participating in on-call rotations
- responding to incidents outside of business hours when needed
Desired Qualifications
- Experience building tooling on top of large language model APIs such as Anthropic Claude or OpenAI
- Background in payment processing or financial services infrastructure
- Familiarity with service management platforms used for tracking incidents, changes, and requests in enterprise engineering environments
- Prior contribution to internal developer platforms or shared engineering tooling
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.