IG Group logo
IG GroupPosted 2 weeks ago

Senior Platform SRE

HybridKraków, Lesser Poland, Poland

Full TimeSenior LevelLarge

Job Summary

Build and own the reliability platform, implementing comprehensive monitoring with OpenTelemetry and distributed tracing to establish SLOs and error budgets. Engineer self-healing capabilities including auto-remediation, error-budget-gated rollbacks, and automated traffic rerouting while designing chaos experiments across the AWS estate. Maintain 24/7 operational readiness through automated deployments, blue/green releases, and zero-downtime patching strategies. Mentor junior SREs and Reliability Champions on production engineering discipline, facilitate blameless post-incident reviews within five days, and evolve SRE standards for the organization. Work with development teams to design SLOs on customer journeys and assist in system design and capacity planning.

Required Qualifications

  • 6+ years of experience in all the below mentioned areas
  • hands-on OpenTelemetry experience (spans, metrics, traces, context propagation) and production use of Honeycomb, Datadog, Dynatrace, or Grafana
  • able to instrument Java or Python services directly
  • proven track record designing customer-meaningful SLIs, setting error budgets, configuring multi-window burn-rate alerts, and working with development teams on reliability measurement
  • experience building pipelines with safety mechanisms: blue/green and canary releases, automated rollback, and DORA metrics integration
  • Kubernetes (EKS, AKS, or GKE) required
  • HashiCorp Nomad is a strong advantage on IG's hybrid estate
  • solid understanding of cloud networking and IaC (Terraform preferred)
  • production-quality coding in Java and/or Python
  • comfortable contributing to application codebases to implement reliability patterns, not just configuring infrastructure around them
  • strong understanding of how large-scale systems fail and how to make them fail safely
  • circuit breakers, bulkheads, idempotency, graceful degradation, and load-shedding
  • high-throughput, low-latency environments preferred
  • on-call experience on production systems
  • blameless PIR facilitation, contributing-factor analysis, and driving action items to closure
  • PagerDuty and ServiceNow familiarity helpful
  • experience designing and executing hypothesis-driven experiments with blast-radius controls and gap-to-impact-tolerance analysis
  • AWS FIS, Gremlin, or equivalent
  • at ease in a guild or community-of-practice model
  • comfortable writing RFCs, presenting at engineering forums, and building standards that others will adopt
  • Track record in high-throughput, production environments (financial services, trading platforms, or similar mission-critical systems preferred)
  • Demonstrated ability to improve system reliability and performance at scale
  • Experience working collaboratively with development teams to implement observability and reliability improvements
  • Strong troubleshooting skills in distributed systems environments
  • Systems thinking approach to problem-solving
  • Excellent communication skills for cross-functional collaboration and technical enablement
  • Ability to balance hands-on development work with operational responsibilities
  • Strong bias toward automation and eliminating manual toil
  • 3 days in the office

Desired Qualifications

  • curious and forward-thinking mindset
  • Java or Python services directly

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce