GE Vernova logo
GE VernovaPosted 2 weeks ago
EXPIRED

SRE Observability SLO Engineer

On-siteMonterrey, Nuevo León, Mexico or Mexico City, Mexico City, Mexico

Full TimeEnterprise

Job Summary

Build and own the full telemetry stack, establishing organization-wide standards for metrics, logs, and distributed traces across all GridOS SaaS services. Define Service Level Indicators and Objectives per product tier, maintain automated compliance reports, and translate these into customer-facing SLAs. Design operational dashboards covering availability, latency, and error rates, while implementing noise-reduced alert policies and escalation routes. Implement synthetic monitoring plans for critical user journeys and API health checks using AWS CloudWatch Synthetics or equivalent tools. After initial v1.0 coverage, transition into an ongoing improvement cadence that expands feature coverage, tunes alert signals, and conducts periodic health reviews to reduce detection and resolution times.

Required Qualifications

  • 2–3 years in SRE, observability engineering, or infrastructure reliability roles
  • Fluent in English
  • Experience with at least one major observability platform — Datadog, Grafana + Prometheus, AWS CloudWatch, Dynatrace, or New Relic
  • Decent understanding of distributed systems telemetry: metrics (Prometheus/CloudWatch), structured logging (CloudWatch Logs Insights, ELK), and distributed tracing (OpenTelemetry, AWS X-Ray)
  • Experience with Kubernetes observability — kube-state-metrics, node exporters, Helm deployed monitoring stacks, and namespace-level resource metrics
  • Proficiency in at least one query/visualization language: PromQL, Splunk SPL, Datadog Query Language, or CloudWatch Logs Insights query syntax
  • Experience enabling monitoring alerts to provide visibility to system health
  • Scripting skills in Python and/or Bash for automation of monitoring configuration and report generation
  • Familiarity with OpenTelemetry (OTel) for vendor-agnostic instrumentation
  • Experience with synthetic monitoring tools — AWS CloudWatch Synthetics, Datadog Synthetics, or Catchpoint
  • Experience in regulated industries — energy, utilities, healthcare — where compliance grade audit trails are required
  • AWS certifications: CloudWatch / Observability specialty, Solutions Architect Associate or Professional
  • Bachelor's Degree in Computer Science

Desired Qualifications

  • Experience with Ansible, Chef or Puppet
  • Experience with AWS and Azure alerting notification
  • Scripting skills in Go, Python, Groovy, Bash
  • Linux Administration Skills

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Find similar roles