Senior SRE / Cloud Engineer
RemoteKyiv, Kyiv City, Ukraine or Ukraine
Job Summary
Operate, monitor, and improve production cloud infrastructure across AWS, Azure, GCP, and hybrid environments while building automation for deployments, health checks, and recovery workflows. Construct monitoring, logging, metrics, tracing, and dashboards to enhance observability across infrastructure, applications, and network dependencies. Participate in incident response, root cause analysis, and post-incident remediation to tune alerts and reduce noise. Partner with engineering teams to define SLOs, SLIs, error budgets, and operational readiness standards. Maintain Infrastructure as Code using Terraform or CloudFormation and create runbooks and escalation procedures. Identify production risks and drive remediation through automation and architecture improvements.
Required Qualifications
- Hands-on production cloud infrastructure experience
- Strong experience with AWS, GCP, Windows, and/or hybrid cloud environments
- Experience building and maintaining observability, monitoring, logging, dashboards, and alerting systems
- Strong troubleshooting skills across infrastructure, networking, application, and cloud service layers
- Experience with Linux systems, networking fundamentals, DNS, TLS, IAM, load balancers, storage, and compute
- Infrastructure as Code experience using Terraform, CloudFormation, Pulumi, or similar tools
- Scripting and automation skills in Bash, Python, Go, or similar languages
- Experience participating in production incident response and postmortem processes
- Familiarity with SRE practices, including SLOs, SLIs, error budgets, toil reduction, and operational readiness
- Ability to work closely with engineering teams to improve reliability and production supportability
- Comfortable troubleshooting live systems
- Improving monitoring coverage
- Building automation
- Partnering with engineering teams to make production systems more reliable
- Operating, monitoring, and improving cloud platforms and services with a focus on uptime, performance, scalability, alerting quality, and operational excellence
- Troubleshooting live systems
- Improving monitoring coverage
- Building automation
- Partnering with engineering teams to make production systems more reliable
- Operating, maintain, and improve production cloud infrastructure across AWS, Azure, GCP, Windows, or hybrid environments
- Build and maintain monitoring, logging, metrics, tracing, dashboards, and alerting for production services
- Improve observability coverage across infrastructure, applications, databases, queues, and network dependencies
- Tune alerts to reduce noise, improve signal quality, and ensure actionable incident response
- Participate in incident response, production troubleshooting, root cause analysis, and post-incident remediation
- Build automation for infrastructure operations, deployments, health checks, runbooks, and recovery workflows
- Partner with engineering teams to define SLOs, SLIs, error budgets, and operational readiness standards
- Support cloud networking, DNS, TLS, load balancing, IAM, storage, compute, and managed service operations
- Improve reliability, availability, performance, and scalability of cloud-hosted systems
- Maintain Infrastructure as Code and configuration management practices for repeatable environments
- Create and maintain runbooks, operational documentation, and escalation procedures
- Identify production risks and drive remediation through automation, architecture improvements, and platform standards
Desired Qualifications
- Kubernetes and Cloud-Native Platforms
- GitOps experience (Flux or Argo CD)
- Jenkins and Ansible
- Service Mesh, Ingress, and API Gateway experience
- High Availability architectures
- Multi-region environments
- Disaster Recovery solutions
- Secrets Management (Vault, AWS Secrets Manager, External Secrets)
- Security, Compliance, and Vulnerability Management
- On-call Operations and Runbook design
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.