Senior DevOps Engineer
$500,000–$500,000 year
RemoteSerbia
Job Summary
Own one of our core platform areas end-to-end: observability (VictoriaMetrics, Grafana, Graylog) or CI/CD (Jenkins scripted pipelines, Harbor, Nexus) — drive its architecture, reliability, and roadmap. Gather requirements, write design docs, decompose tasks, implement, and deliver to production while owning operational health. Support developers by deploying and monitoring applications on Kubernetes, troubleshooting builds, and participating in chat duty. Investigate production incidents as the senior escalation point, lead post-mortems, and implement systemic fixes. Mentor less experienced engineers through design discussions and reviews. Use AI in all aspects of day-to-day work. Participate in on-call rotations and raise the bar for how on-call works.
Required Qualifications
- 6+ years as a DevOps Engineer / SRE (or very close responsibilities)
- Track record of owning technical initiatives end-to-end — from requirements and technical design through production delivery
- Confident Linux skills (we use Ubuntu)
- Working knowledge of the Prometheus stack: metric types, exporters, and how alerting works — enough to navigate and extend an existing setup
- Hands-on experience with CI/CD: pipeline design, build orchestration, artifact delivery
- Containers: Docker, image building, registries
- Ansible
- Git
- Experience with Bash or Python scripting for automation and observability (writing exporters, eliminating routine work)
- Production/on-call experience: diagnosing incidents, restoring service, leading post-mortems
- Experience mentoring less experienced engineers
- Ownership and attention to detail
- Must have solid hands-on experience in two or more of the areas below: VictoriaMetrics / Prometheus stack at scale: architecture, cardinality control, exporters, alerting infrastructure
- Must have solid hands-on experience in two or more of the areas below: Log pipelines at scale: Graylog / VictoriaLogs / ELK — collection (fluent bit or similar), retention, sharding, performance
- Must have solid hands-on experience in two or more of the areas below: Jenkins scripted pipelines: shared libraries, pipeline infrastructure, build agent fleets
- Must have solid hands-on experience in two or more of the areas below: Container registries and artifact management: Harbor, Nexus, base images, image policies
- Must have solid hands-on experience in two or more of the areas below: Operating applications on Kubernetes: Helm, workload monitoring and log delivery, deploy troubleshooting
- Must have solid hands-on experience in two or more of the areas below: Grafana: dashboards as code, alerting, performance at scale
- Must be able to lift 50 lbs
Desired Qualifications
- Great if you've worked with any of the following: Analytics & DS platforms: JupyterHub, Airflow, Tableau, MLflow, Airbyte — deployment, maintenance, resource limits. Building platform around these tools to improve Quality of Life for Analytics
- Great if you've worked with any of the following: Remote development environments and AI agent execution environments. E.g. Coder/Telepresence
- Great if you've worked with any of the following: Bare-metal Kubernetes: provisioning, networking, scaling
- Great if you've worked with any of the following: Flux and GitOps
- Great if you've worked with any of the following: Terraform
- Great if you've worked with any of the following: Sentry on-premise: operating self-hosted error tracking
- Great if you've worked with any of the following: ClickHouse, MongoDB
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.