Reward Gateway logo
Reward GatewayPosted 2 months ago

Senior Software Engineer

$70,000–$85,000 year

HybridLondon, England, United Kingdom

Full TimeSenior LevelMedium

Job Summary

Provide high-quality L2.5 support for PHP applications on EKS with MySQL, executing scripted remediation, configuration changes, and safe code-level fixes within defined guardrails. Automate repetitive tasks and maintain runbooks to reduce manual toil and improve MTTR. Act as a first responder for application incidents, triaging, diagnosing, and coordinating remediation using Datadog, Kibana, and Heap for impact analysis. Contribute to service onboarding reviews, post-incident learning, and operational KPI reporting. Work in a hybrid model with a sustainable on-call rotation and collaborate with product teams to enhance operational capabilities.

Required Qualifications

  • Proven experience in application support or operations engineering in cloud environments, ideally supporting PHP services running on Kubernetes (EKS) with MySQL backends
  • Hands-on capability in at least one backend language (PHP preferred; Python or similar also valuable) sufficient to read, diagnose, and write safe operational scripts and minor fixes under guardrails
  • Practical Kubernetes skills for operations: kubectl/Helm basics, investigating pods/deployments, reading logs/events, understanding readiness/liveness probes, and performing safe rollouts/rollbacks within documented guardrails
  • MySQL operational fluency: connection and pool issues, slow query detection, query plan basics, common remediation patterns (e.g., indexing recommendations to hand to L3, safe data fixes under runbook guardrails), and understanding of replication/backup implications
  • Strong experience using Datadog (APM/metrics/traces/dashboards/alerts) for investigation and detection; confident using Kibana for log exploration and correlation; ability to leverage Heap to assess user impact and prioritize remediation
  • Familiarity with ITSM tooling (e.g., Jira Service Management) and ITIL-aligned incident and problem management processes
  • Strong communication skills; clear, concise documentation; collaborative approach focused on reducing toil, increasing automation, and raising the quality bar
  • This role offers a hybrid work model to be present in our London office twice a week
  • Standard hours are 9am - 6pm, Mon - Fri
  • 1 day in every 4 is on call, paid at 1.5x hourly rate
  • On call hours are 6pm - 9am
  • If your on call day falls on a weekend, 24 hour on call cover is required

Desired Qualifications

  • Python
  • Experience with feature flag platforms and configuration-as-code within safe operational guardrails
  • Familiarity with AWS services that commonly interface with PHP/EKS workloads (e.g., CloudWatch, ALB, S3, SQS) and how they surface in Datadog and Kibana
  • Exposure to service onboarding/operability reviews, SLOs, and contributing to a Service Catalogue
  • Experience balancing incident response with proactive improvement work in Agile contexts; strong documentation discipline

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce