Software Engineer 3, Platform
$88,540–$88,540 year
HybridChicago, Illinois, United States or Atlanta, Georgia, United States
Job Summary
Own and evolve meaningful pieces of the multi-region Kubernetes, observability, and CI/CD infrastructure systems. Manage Terraform infrastructure-as-code across AWS accounts, build GitLab CI/CD pipelines, and enforce platform standards like RBAC and admission webhooks. Respond to and drive resolution of production platform incidents, focusing on learning and system improvement. Act as a reviewer for infrastructure changes including Terraform, Kubernetes configs, and observability settings. Turn recurring access requests into self-service workflows while maintaining deep operational depth in systems like Prometheus, Alertmanager, and AWS networking.
Required Qualifications
- You must be work authorized in the United States without the need for employer sponsorship
- This is a hybrid role requiring 3 days a week in office
- 3+ years of experience in software or infrastructure engineering
- Bachelor's degree or equivalent experience
- Hands-on production experience with Kubernetes and at least one major cloud (AWS preferred)
- Real operational depth in at least one system we own beyond the cluster - most valuably the observability stack (Prometheus/Alertmanager at scale), but AWS networking, Vault, or artifact/CI infrastructure also count
- Comfortable owning infrastructure-as-code (Terraform) and CI/CD pipelines
- Can reason about tradeoffs and communicate the pros and cons of multiple approaches
- Effective communication; thrives in a collaborative, pair-friendly team culture
- Deep Prometheus and Alertmanager knowledge
- Multi-region EKS clusters: upgrades, node group and Karpenter management, controller lifecycle, and add-on / configuration management
- VPC and subnet design, CIDR management, VPC peering, Route53, security groups, and NAT gateway topology across accounts and regions
- GitLab administration (runner fleet, cache, access - not just pipeline authoring)
- GitOps delivery through ArgoCD
- Nexus artifact repository including its storage lifecycle
- Vault secrets management
- IAM roles and service accounts for apps in clusters
- Cluster permission management for audit compliance
- AI model access management
- OpenCost, EBS orphan cleanup, cost anomaly investigation, and rightsizing attribution
- Terraform
- AWS (IAM, EKS, S3, EBS)
- ArgoCD
- GitLab CI/CD
- Nexus (artifact registry)
- Docker
- container image build pipelines
- Kubernetes controllers/operators (reconciliation patterns, restart safety)
Desired Qualifications
- Go experience is a plus, not required
- AWS networking depth (Transit Gateway, multi-account topology)
- Prometheus long-term storage / sharding (Thanos, Cortex, Mimir, or equivalent)
- Policy-as-code (Kyverno / OPA)
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.