Senior AI Operations Engineer
HybridBarcelona, Catalonia, Spain
Job Summary
Design, build, and maintain the centralized observability layer using Datadog, New Relic, or Grafana while architecting AI-augmented tools for runbook querying and incident pattern analysis. Lead complex incident response, perform root-cause analysis, and drive the continuous improvement flywheel by automating toil and eliminating reactive work. Mentor L1 operators and junior L2 engineers, preparing flywheel metrics and fostering a culture where operational excellence is an engineering discipline. Contribute patches to product engineering codebases and implement the Operational Readiness Gate to ensure platform operability before handover. Work within a three-tier structure across Azure and AWS, balancing deep technical execution with collaborative stakeholder engagement to shape the future of AI-driven operations.
Required Qualifications
- BSc/MSc degree in Computer Science or related quantitative or analytical field
- Significant hands-on experience as a Site Reliability Engineer or platform operations engineer at scale
- Strong expertise in observability platforms (Datadog, New Relic, Grafana, Splunk, or equivalent) including dashboard design, alerting strategies, and telemetry pipeline implementation
- Strong working knowledge of OpenTelemetry, distributed tracing, and structured logging standards
- Proven track record of designing and implementing automation that materially reduces operational toil — strong scripting skills in Python and Bash, with the ability to build robust tooling beyond one-off scripts
- Experience assessing platform readiness and contributing to operational handover processes
- Strong hands-on skills with cloud infrastructure (Azure and/or AWS) including container orchestration, serverless architectures, and managed services
- Experience running or contributing significantly to post-mortem processes and translating findings into preventive engineering work
- Ability to implement precise technical solutions — you can instrument a system, configure meaningful alerts, and build automation that intervenes at the right point
- Experience mentoring junior engineers and contributing to team development
- Experience applying AI/ML to operational challenges — intelligent alerting, automated diagnosis, predictive incident detection, or conversational operations interfaces
- Familiarity with the AstraZeneca technology estate or regulated pharmaceutical environments
- Experience operating platforms that serve AI/ML workloads (LLM inference, model serving, data pipelines)
- ITIL, SRE, or operational excellence certifications or equivalent practical frameworks
- Experience working within multi-tier support structures with clear escalation paths
- Infrastructure-as-code expertise (Terraform, CloudFormation)
- Awareness to GxP and audit trail awareness
- Minimum of three days per week from the office
Desired Qualifications
- Experience applying AI/ML to operational challenges — intelligent alerting, automated diagnosis, predictive incident detection, or conversational operations interfaces
- Familiarity with the AstraZeneca technology estate or regulated pharmaceutical environments
- Experience operating platforms that serve AI/ML workloads (LLM inference, model serving, data pipelines)
- ITIL, SRE, or operational excellence certifications or equivalent practical frameworks
- Experience working within multi-tier support structures with clear escalation paths
- Infrastructure-as-code expertise (Terraform, CloudFormation)
- Awareness to GxP and audit trail awareness
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.