PostHog logo
PostHogPosted 3 weeks ago

Site Reliability Engineer (US - Central/Eastern time)

Remote

Full TimeMediumTECH

Job Summary

Operate EKS clusters across multiple environments using Karpenter autoscaling, Cilium networking, and ArgoCD-driven GitOps deployments. Manage AWS account organization, provisioning, networking, and access control while maintaining Terraform/Terragrunt infrastructure as code platforms. Reduce operational stress by designing safe automation for traffic-heavy workloads and building self-healing tools for deploys, backups, and incident response. Participate in on-call rotation to ensure production reliability across petabytes of data and thousands of cores. Optimize cloud spend and eliminate repeat pain points through code. Join a remote team focused on shipping fast, autonomous product development for over 450,000 organizations.

Required Qualifications

  • Deep hands-on experience with Kubernetes in production (EKS preferred)
  • Strong experience operating production infrastructure on AWS
  • Experience automating infrastructure using Terraform or Terragrunt at scale
  • Solid understanding of Linux systems
  • Experience supporting stateful systems
  • Ability to debug and reason about performance and reliability issues in production
  • Comfortable owning systems end-to-end, including on-call responsibilities

Desired Qualifications

  • Experience with GitOps workflows (ArgoCD) and CI/CD pipelines (GitHub Actions)
  • Experience with building AI agent-enabled base-level infra services for teams that move fast
  • Familiarity with multi-region infrastructure and the consistency/availability tradeoffs that come with it

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce