Impact logo
ImpactPosted 1 month ago

Site Reliability Engineer

On-siteCape Town, Western Cape, South Africa

Full TimeMediumPartnership Management Software

Job Summary

Define service-level objectives and indicators for GCP services, establishing error budgets and baselines for critical business processes. Build and standardize observability using Open Telemetry, Grafana, and Prometheus to surface system health without leaking PII, while defining actionable alerting and runbook conventions. Own root-cause analysis practices by driving investigations toward durable fixes and preventative actions, partnering with squads to improve troubleshooting habits. Manage vulnerability postures across JVM, Spring dependencies, and container images, ensuring secrets hygiene and least-privilege access alongside forensic-replay history for compliance. Troubleshoot issues across the stack including JVM tuning, database optimization, and CI/CD hardening, while analyzing capacity and costs to inform efficiency improvements. This SRE role prioritizes system stability and data integrity over feature velocity within the Content Intelligence & Regulatory Apps Group.

Required Qualifications

  • 3+ years in SRE, software engineering, or systems/operations roles supporting production services
  • Solid understanding of systems and application design
  • Able to read, debug, and make changes to Java/Spring code today
  • Willingness and aptitude to deepen JVM expertise (garbage collection, memory, thread pools) on the job
  • Experience operating services on a major cloud (ideally GCP/GKE)
  • Exposure to Kubernetes and containers
  • Proficiency with metrics, logging, tracing/APM, and alerting
  • Experience with relational database monitoring and SQL tuning (MySQL preferred)
  • Experience with basic indexing and query optimization
  • Proficiency in shell scripting
  • Comfort automating routine operational tasks
  • Practical understanding of operational security for production systems
  • Understanding of secrets management, least-privilege access, and handling sensitive data safely in logs and telemetry
  • Familiarity with defining or operating against SLOs/SLIs
  • Strong desire and aptitude to build that practice
  • Collaborative working style with the ability to influence and enable other engineers
  • Ability to prioritize, work independently, and focus on simple, efficient, and reliable solutions
  • B.S. in Computer Science or a related field
  • Equivalent practical experience

Desired Qualifications

  • Experience with the Spring/Spring Boot ecosystem and JVM application servers
  • Experience with tools like Quartz, Temporal, or similar
  • Familiarity with Pub/Sub or Kafka and service-to-service integration (REST/Feign, gRPC)
  • Experience with Prometheus/PromQL, Grafana, and log analytics
  • Familiarity with using AI tools to assist in engineering workflows
  • Familiarity with secrets management (Vault, GCP Secret Manager, SOPS)
  • Familiarity with code-quality gates (SonarQube)
  • Familiarity with CI/CD + GitOps (Jenkins, ArgoCD/Helm)

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce