Site Reliability Engineer (US - Central/Eastern time)
Remote
Job Summary
Operate EKS clusters across multiple environments using Karpenter autoscaling, Cilium networking, and ArgoCD-driven GitOps deployments. Manage AWS account organization, provisioning, networking, and access control while maintaining Terraform/Terragrunt infrastructure as code platforms. Reduce operational stress by designing safe automation for traffic-heavy workloads and building self-healing tools for deploys, backups, and incident response. Participate in on-call rotation to ensure production reliability across petabytes of data and thousands of cores. Optimize cloud spend and eliminate repeat pain points through code. Join a remote team focused on shipping fast, autonomous product development for over 450,000 organizations.
Required Qualifications
- Deep hands-on experience with Kubernetes in production (EKS preferred)
- Strong experience operating production infrastructure on AWS
- Experience automating infrastructure using Terraform or Terragrunt at scale
- Solid understanding of Linux systems
- Experience supporting stateful systems
- Ability to debug and reason about performance and reliability issues in production
- Comfortable owning systems end-to-end, including on-call responsibilities
Desired Qualifications
- Experience with GitOps workflows (ArgoCD) and CI/CD pipelines (GitHub Actions)
- Experience with building AI agent-enabled base-level infra services for teams that move fast
- Familiarity with multi-region infrastructure and the consistency/availability tradeoffs that come with it
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.