Principal Platform Engineer
On-siteLondon, England, United Kingdom
Job Summary
Design, build, and operate global trading platform infrastructure across AWS and Azure, managing core networking, Kubernetes, and high-availability architectures. Lead the adoption of AI-driven operations by prototyping and productionizing workflows for incident triage, runbook automation, and log analysis. Establish SLOs, error budgets, and observability standards while reducing toil through automation and scripting. Mentor senior engineers and act as design authority for cross-team projects, ensuring new services are resilient and operable from day one. Balance reliability investment against delivery on the technical roadmap.
Required Qualifications
- A background in Operations or SRE running highly available, redundant production platforms
- Deep hands-on experience with AWS (VPC design, networking, EKS)
- Deep hands-on experience with Azure (AKS, Key Vault, networking)
- Genuinely multi-cloud experience (not one cloud plus a certification)
- Strong Networking fundamentals: TCP/IP, routing concepts, firewalls, load balancing, hybrid connectivity (Direct Connect / ExpressRoute, VPNs)
- Production Kubernetes experience at scale, including day-2 operations (upgrades, capacity, security, multi-cluster)
- Infrastructure-as-code and automation as a default working style (Terraform, Ansible, or similar)
- Strong scripting in Python, Go, or Bash
- Demonstrable, practical use of AI tooling to improve engineering or operational workflows
- The credibility and communication skills to mentor senior engineers and influence design decisions without formal authority
- Understanding of different database technologies
Desired Qualifications
- Experience in trading, exchanges, market data, fintech, or another latency- and availability-sensitive domain
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.