Senior AI Operations Engineer
HybridBarcelona, Catalonia, Spain
Job Summary
Design, build, and maintain the centralized observability layer using Datadog, New Relic, or Splunk while architecting AI-augmented tools for runbook querying and incident pattern analysis. Lead complex incident response, perform root-cause analysis, and produce actionable post-mortems, contributing patches to product engineering codebases to address underlying issues. Build and maintain automation that reduces toil to under 50% of SRE time, developing runbooks and instrumentation for operational readiness gates before platform handover. Mentor L1 operators and junior L2 engineers, fostering a culture where engineers drive operational excellence through systems thinking and continuous improvement. Work within a three-tier operations structure to shape the future of AstraZeneca's digital and data-led enterprise.
Required Qualifications
- BSc/MSc degree in Computer Science or related quantitative or analytical field
- Significant hands-on experience as a Site Reliability Engineer or platform operations engineer at scale
- Strong expertise in observability platforms (Datadog, New Relic, Grafana, Splunk, or equivalent) including dashboard design, alerting strategies, and telemetry pipeline implementation
- Strong working knowledge of OpenTelemetry, distributed tracing, and structured logging standards
- Proven track record of designing and implementing automation that materially reduces operational toil — strong scripting skills in Python and Bash, with the ability to build robust tooling beyond one-off scripts
- Experience assessing platform readiness and contributing to operational handover processes
- Strong hands-on skills with cloud infrastructure (Azure and/or AWS) including container orchestration, serverless architectures, and managed services
- Experience running or contributing significantly to post-mortem processes and translating findings into preventive engineering work
- Ability to implement precise technical solutions — you can instrument a system, configure meaningful alerts, and build automation that intervenes at the right point
- Experience mentoring junior engineers and contributing to team development
- Experience applying AI/ML to operational challenges — intelligent alerting, automated diagnosis, predictive incident detection, or conversational operations interfaces
- Familiarity with the AstraZeneca technology estate or regulated pharmaceutical environments
- Experience operating platforms that serve AI/ML workloads (LLM inference, model serving, data pipelines)
- ITIL, SRE, or operational excellence certifications or equivalent practical frameworks
- Experience working within multi-tier support structures with clear escalation paths
- Infrastructure-as-code expertise (Terraform, CloudFormation)
- Awareness to GxP and audit trail awareness
- Minimum of three days per week from the office
Desired Qualifications
- Experience applying AI/ML to operational challenges — intelligent alerting, automated diagnosis, predictive incident detection, or conversational operations interfaces
- Familiarity with the AstraZeneca technology estate or regulated pharmaceutical environments
- Experience operating platforms that serve AI/ML workloads (LLM inference, model serving, data pipelines)
- ITIL, SRE, or operational excellence certifications or equivalent practical frameworks
- Experience working within multi-tier support structures with clear escalation paths
- Infrastructure-as-code expertise (Terraform, CloudFormation)
- Awareness to GxP and audit trail awareness
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.