Lead Site Reliability Engineer
On-siteLondon, England, United Kingdom
Job Summary
Engage daily with traders across Equities, Fixed Income, and FX to understand workflows and reliability priorities while acting as a trusted engineering partner for desk stability. Support live trading environments through incident response, root cause analysis, and post mortem leadership, contributing directly to the codebase in Java, Kotlin, and Python for performance optimizations and automation. Lead the design and rollout of modern SRE patterns including automated remediation and self-healing workflows, utilizing enterprise-authorized AI capabilities to accelerate incident triage and validate operational data. Build monitoring, alerting, and distributed tracing tooling across global environments while collaborating with infrastructure, networking, and cybersecurity teams to ensure end-to-end reliability. Operate within a globally distributed organization in EMEA, US, and APAC, driving improvements in latency, throughput, and stability for high-volume trading applications.
Required Qualifications
- Strong hands on experience in front office trading environments or similarly high pressure, low latency domains
- Proficiency with SRE tooling and techniques, including FIX messaging, Kafka, Grafana, Splunk, ITRS Geneos, Dynatrace, InfluxDB, MQ (IBM MQ or similar), Oracle DB
- Demonstrated experience using enterprise-authorized AI capabilities within the work environment to improve SRE workflows (e.g., incident investigation support and knowledge capture) with strong validation habits and awareness of data sensitivity
- Ability to evaluate AI-assisted operational recommendations for correctness and risk, define appropriate guardrails for team usage, and ensure outcomes align to resiliency and security expectations
- Deep knowledge of reliability engineering principles: SLIs/SLOs, real-time telemetry, disaster recovery planning, capacity planning, and performance tuning
- Experience designing and implementing observability frameworks for mission critical systems
- Proven ability to lead incident response and drive long term remediation
- Solid programming skills in Python, Java, or Kotlin, with the ability to contribute production grade code
- Experience with microservices, distributed systems, and event driven architectures
- Strong understanding of CI/CD pipelines, automated testing, and deployment strategies
- Comfortable interacting directly with traders and senior stakeholders
- Excellent communication skills, especially when translating technical issues into business impact
- Ability to operate calmly and decisively in high pressure situations
- Strong leadership presence with a collaborative mindset
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.