SRE Observability SLO Engineer
On-siteMonterrey, Nuevo León, Mexico or Mexico City, Mexico City, Mexico
Job Summary
Build and own the full telemetry stack for GridOS SaaS services, establishing instrumentation standards, SLO dashboards, and synthetic monitors to ensure real-time reliability confidence. Implement metrics collection for Kubernetes-hosted services and define data retention policies to maintain economic sustainability. Partner with product engineering and customer stakeholders to define SLIs, SLOs, and SLAs, while governing the SLO review cycle and facilitating monthly reliability assessments. Design operational dashboards covering the Golden Signals and implement alert policies with noise-reduction practices, including symptom-based alerting and multi-window burn-rate rules. Establish alert routing, escalation policies, and on-call schedules, then transition into a roadmap-aligned improvement cycle to expand coverage, tune signal-to-noise, and retire stale monitors.
Required Qualifications
- 2–3 years in SRE, observability engineering, or infrastructure reliability roles
- Fluent in English
- Experience with at least one major observability platform — Datadog, Grafana + Prometheus, AWS CloudWatch, Dynatrace, or New Relic
- Decent understanding of distributed systems telemetry: metrics (Prometheus/CloudWatch), structured logging (CloudWatch Logs Insights, ELK), and distributed tracing (OpenTelemetry, AWS X-Ray)
- Experience with Kubernetes observability — kube-state-metrics, node exporters, Helm deployed monitoring stacks, and namespace-level resource metrics
- Proficiency in at least one query/visualization language: PromQL, Splunk SPL, Datadog Query Language, or CloudWatch Logs Insights query syntax
- Experience enabling monitoring alerts to provide visibility to system health
- Scripting skills in Python and/or Bash for automation of monitoring configuration and report generation
- Familiarity with OpenTelemetry (OTel) for vendor-agnostic instrumentation
- Experience with synthetic monitoring tools — AWS CloudWatch Synthetics, Datadog Synthetics, or Catchpoint
- Experience in regulated industries — energy, utilities, healthcare — where compliance grade audit trails are required
- AWS certifications: CloudWatch / Observability specialty, Solutions Architect Associate or Professional
- Bachelor's Degree in Computer Science
Desired Qualifications
- Experience with Cloud Technologies - AWS Cloud Infrastructure - EKS, RDS, MSK, S3, EC2, EBS, SQS, etc.
- Experience with Kubernetes - EKS, Rancher Deployment and Configuration Tools - Ansible, Chef or Puppet
- Experience with Observability tools and technology - Datadog, Splunk, NewRelic, etc.
- Experience with Alerting and notification - AWS and Azure alerting notification
- Scripting - Go, Python, Groovy, Bash
- Linux Administration Skills
- Critical thinker; able to quickly adapt to changing environments
- A hacker or tinkerer at heart
- Risk taker, not afraid to think outside the box or challenge the status quo
- Emotional Intelligence, ability to influence up and out and the ability to work independently
- Must be a team player with a strong desire to win
- Passionate about continuously learning
- Highly organized and efficient; able to balance competing priorities and execute accordingly
- Strong oral and written communication skills
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.