Senior DevOps Engineer
$500,000–$500,000 year
RemoteGeorgia, United States
Job Summary
Own core platform areas end-to-end: observability (VictoriaMetrics, Grafana, Graylog) or CI/CD (Jenkins, Harbor, Nexus), driving architecture, reliability, and roadmap execution. Lead technical initiatives from requirements and design through production rollout and operational health management. Support developers by deploying and monitoring applications on Kubernetes and on-premise servers, troubleshooting builds, and participating in developer support chat duty. Investigate production incidents as the senior escalation point, leading resolution and post-mortems while raising on-call standards. Mentor less experienced engineers through design discussions and code reviews. Automate repetitive operations and provisioning to eliminate toil. Use AI for research, troubleshooting, and development tasks.
Required Qualifications
- 6+ years as a DevOps Engineer / SRE (or very close responsibilities)
- Track record of owning technical initiatives end-to-end — from requirements and technical design through production delivery
- Confident Linux skills (we use Ubuntu)
- Working knowledge of the Prometheus stack: metric types, exporters, and how alerting works — enough to navigate and extend an existing setup
- Hands-on experience with CI/CD: pipeline design, build orchestration, artifact delivery
- Containers: Docker, image building, registries
- Ansible
- Git
- Experience with Bash or Python scripting for automation and observability (writing exporters, eliminating routine work)
- Production/on-call experience: diagnosing incidents, restoring service, leading post-mortems
- Experience mentoring less experienced engineers
- Ownership and attention to detail
- Must have solid hands-on experience in two or more of the areas below: VictoriaMetrics / Prometheus stack at scale: architecture, cardinality control, exporters, alerting infrastructure
- Must have solid hands-on experience in two or more of the areas below: Log pipelines at scale: Graylog / VictoriaLogs / ELK — collection (fluent bit or similar), retention, sharding, performance
- Must have solid hands-on experience in two or more of the areas below: Jenkins scripted pipelines: shared libraries, pipeline infrastructure, build agent fleets
- Must have solid hands-on experience in two or more of the areas below: Container registries and artifact management: Harbor, Nexus, base images, image policies
- Must have solid hands-on experience in two or more of the areas below: Operating applications on Kubernetes: Helm, workload monitoring and log delivery, deploy troubleshooting
- Must have solid hands-on experience in two or more of the areas below: Grafana: dashboards as code, alerting, performance at scale
- Must be able to effectively lead and coach others
- Must be available for weekend shifts
- Valid driver's license
- Must be able to lift 50 lbs
- Must be able to effectively lead and coach others
- Must be available for weekend shifts
- Valid driver's license
- Must be able to lift 50 lbs
Desired Qualifications
- Great if you've worked with any of the following: Analytics & DS platforms: JupyterHub, Airflow, Tableau, MLflow, Airbyte — deployment, maintenance, resource limits
- Great if you've worked with any of the following: Remote development environments and AI agent execution environments. E.g. Coder/Telepresence
- Great if you've worked with any of the following: Bare-metal Kubernetes: provisioning, networking, scaling
- Great if you've worked with any of the following: Flux and GitOps
- Great if you've worked with any of the following: Terraform
- Great if you've worked with any of the following: Sentry on-premise: operating self-hosted error tracking
- Great if you've worked with any of the following: ClickHouse, MongoDB
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.