Jalasoft logo
JalasoftPosted 2 weeks ago

Site Reliability Engineer - Azure, Observability and Scripting

RemoteArgentina or Colombia

Full TimeLarge

Job Summary

Conduct site reliability engineering for cloud-native platforms on Microsoft Azure and Kubernetes, focusing on observability, automation, and operational excellence. Design and implement Azure Monitor, Log Analytics, Prometheus, and Grafana solutions, including workspace configuration, alerting rules, and dashboard creation. Manage incident response, runbook authoring, and on-call practices while defining service level objectives and error budgets. Execute backup, restore, and disaster recovery testing against RPO and RTO targets, and troubleshoot Kubernetes workloads, clusters, and upgrades. Modify infrastructure as code using Terraform or Bicep and automate pipelines with Azure DevOps. This role supports Jalasoft's mission to ensure system availability and performance.

Required Qualifications

  • 6+ years of experience
  • 3+ years of experience operating Kubernetes in production
  • Site reliability engineering or production operations for Kubernetes workloads at scale
  • Azure Monitor, Log Analytics and KQL, including workspace design, data collection rules and retention strategy
  • Prometheus and Grafana: metrics and exporters, recording and alerting rules, and dashboard design
  • Definition and implementation of service level indicators, objectives and error budgets
  • Alerting and incident response design, including runbook authoring and on-call practice
  • Backup, restore and disaster recovery design and testing, including validation against RPO and RTO targets
  • Kubernetes operations: workload troubleshooting, resource management and cluster upgrades
  • Ability to read and modify infrastructure as code (Terraform or Bicep) and Azure DevOps pipelines
  • Scripting in Python, PowerShell or Bash
  • Professional working English

Desired Qualifications

  • Azure Managed Prometheus and Azure Managed Grafana
  • OpenTelemetry instrumentation and distributed tracing
  • Azure Backup, Azure Site Recovery, and snapshot-based recovery of virtual machines
  • Chaos engineering or structured game day practice
  • Database-layer observability, particularly for Oracle
  • Cost and capacity management for AKS estates
  • Incident management tooling and postmortem practice
  • Certification: CKA, AZ-400 or equivalent

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce