Paycom logo
PaycomPosted 3 weeks ago

Site Reliability Engineer (SRE)

$100,000–$115,000 year

On-siteAtlanta, Georgia, United States

Full Time

Job Summary

Provide 24x7 production support through an on-call rotation to ensure application availability and rapid incident response. Monitor production applications and infrastructure using Datadog, Splunk, and CloudWatch to identify performance bottlenecks and operational issues. Troubleshoot Java-based applications running on Amazon EKS, executing deployments, hotfixes, and rollbacks while performing root cause analysis to prevent recurring incidents. Develop automation scripts in Python or Bash to eliminate repetitive tasks and build self-healing processes that improve system reliability. Partner with Development teams to enhance application resiliency and scalability, leveraging AI-powered tools to optimize monitoring and operational efficiency.

Required Qualifications

  • 5+ years of experience supporting production applications in an SRE, DevOps, Production Support, or Site Reliability Engineering role
  • Strong experience supporting Java-based applications in production environments
  • Hands-on experience with AWS services including EKS, EC2, ALB/NLB, RDS, IAM, Route 53, CloudWatch, S3, and VPC
  • Experience with Kubernetes (Amazon EKS), Docker, and containerized application deployments
  • Strong experience using Datadog / Splunk for infrastructure monitoring, APM, troubleshooting, dashboards, alerting, and log analysis
  • Experience performing production deployments through CI/CD pipelines (Jenkins, GitHub Actions, Argo CD, or similar)
  • Experience supporting MySQL and Oracle databases from an application support perspective
  • Proficiency in Python, Bash, or other scripting languages for automation
  • Strong Linux system administration and troubleshooting skills
  • Excellent troubleshooting skills across distributed applications, networking, and cloud infrastructure
  • Knowledge of networking fundamentals including DNS, TCP/IP, HTTP/HTTPS, TLS, load balancing, and firewalls
  • Experience with incident management, problem management, and change management processes
  • Experience using AI-assisted development tools or agentic AI systems to improve operational efficiency

Desired Qualifications

  • Experience with Helm and GitOps deployment models
  • Experience with Terraform or Infrastructure as Code
  • Familiarity with Prometheus, Grafana, or OpenTelemetry
  • Experience supporting microservices architectures
  • Knowledge of JVM tuning and Java performance optimization
  • Experience with AWS Auto Scaling, Karpenter, or Cluster Autoscaler

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce