Senior DevOps Engineer
$62,400–$70,800 year
RemoteArmenia, South Carolina, United States
Job Summary
Own core platform areas end-to-end: observability (VictoriaMetrics, Grafana, Graylog) or CI/CD (Jenkins, Harbor, Nexus), driving architecture, reliability, and roadmap from requirements through production delivery. Lead technical design, implement solutions, and maintain operational health while automating repetitive tasks and codifying provisioning. Act as the senior escalation point for complex incidents, investigating production outages, leading post-mortems, and implementing systemic fixes. Mentor less experienced engineers through design discussions, code reviews, and pairing sessions to catch debt-inducing shortcuts. Support developers with deployment troubleshooting, metrics, alerts, and logs across on-premise servers and Kubernetes environments. Participate in on-call rotations and raise the bar for incident response. Use AI for researching, troubleshooting, and developing.
Required Qualifications
- 6+ years as a DevOps Engineer / SRE (or very close responsibilities)
- Track record of owning technical initiatives end-to-end — from requirements and technical design through production delivery
- Confident Linux skills (we use Ubuntu)
- Working knowledge of the Prometheus stack: metric types, exporters, and how alerting works — enough to navigate and extend an existing setup
- Hands-on experience with CI/CD: pipeline design, build orchestration, artifact delivery
- Containers: Docker, image building, registries
- Ansible
- Git
- Experience with Bash or Python scripting for automation and observability (writing exporters, eliminating routine work)
- Production/on-call experience: diagnosing incidents, restoring service, leading post-mortems
- Experience mentoring less experienced engineers
- Ownership and attention to detail
- Must have solid hands-on experience in two or more of the areas below: VictoriaMetrics / Prometheus stack at scale: architecture, cardinality control, exporters, alerting infrastructure
- Must have solid hands-on experience in two or more of the areas below: Log pipelines at scale: Graylog / VictoriaLogs / ELK — collection (fluent bit or similar), retention, sharding, performance
- Must have solid hands-on experience in two or more of the areas below: Jenkins scripted pipelines: shared libraries, pipeline infrastructure, build agent fleets
- Must have solid hands-on experience in two or more of the areas below: Container registries and artifact management: Harbor, Nexus, base images, image policies
- Must have solid hands-on experience in two or more of the areas below: Operating applications on Kubernetes: Helm, workload monitoring and log delivery, deploy troubleshooting
- Must have solid hands-on experience in two or more of the areas below: Grafana: dashboards as code, alerting, performance at scale
- Must be able to effectively lead and coach others
- Must be available for weekend shifts
- Valid driver's license
- Must be able to lift 50 lbs
- Must be able to effectively lead and coach others
- Must be available for weekend shifts
- Valid driver's license
- Must be able to lift 50 lbs
Desired Qualifications
- Great if you've worked with any of the following: Analytics & DS platforms: JupyterHub, Airflow, Tableau, MLflow, Airbyte — deployment, maintenance, resource limits
- Great if you've worked with any of the following: Remote development environments and AI agent execution environments. E.g. Coder/Telepresence
- Great if you've worked with any of the following: Bare-metal Kubernetes: provisioning, networking, scaling
- Great if you've worked with any of the following: Flux and GitOps
- Great if you've worked with any of the following: Terraform
- Great if you've worked with any of the following: Sentry on-premise: operating self-hosted error tracking
- Great if you've worked with any of the following: ClickHouse, MongoDB
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.