AstraZeneca logo
AstraZenecaPosted 1 week ago

Senior AI Operations Engineer

HybridBarcelona, Catalonia, Spain

Full TimeSenior LevelEnterprise

Job Summary

Design, build, and maintain the centralized observability layer using Datadog, New Relic, or Grafana while architecting AI-augmented tools for runbook querying and incident pattern analysis. Lead complex incident response, perform root-cause analysis, and drive the continuous improvement flywheel by automating toil and eliminating reactive work. Mentor L1 operators and junior L2 engineers, preparing flywheel metrics and fostering a culture where operational excellence is an engineering discipline. Contribute patches to product engineering codebases and implement the Operational Readiness Gate to ensure platform operability before handover. Work within a three-tier structure across Azure and AWS, balancing deep technical execution with collaborative stakeholder engagement to shape the future of AI-driven operations.

Required Qualifications

  • BSc/MSc degree in Computer Science or related quantitative or analytical field
  • Significant hands-on experience as a Site Reliability Engineer or platform operations engineer at scale
  • Strong expertise in observability platforms (Datadog, New Relic, Grafana, Splunk, or equivalent) including dashboard design, alerting strategies, and telemetry pipeline implementation
  • Strong working knowledge of OpenTelemetry, distributed tracing, and structured logging standards
  • Proven track record of designing and implementing automation that materially reduces operational toil — strong scripting skills in Python and Bash, with the ability to build robust tooling beyond one-off scripts
  • Experience assessing platform readiness and contributing to operational handover processes
  • Strong hands-on skills with cloud infrastructure (Azure and/or AWS) including container orchestration, serverless architectures, and managed services
  • Experience running or contributing significantly to post-mortem processes and translating findings into preventive engineering work
  • Ability to implement precise technical solutions — you can instrument a system, configure meaningful alerts, and build automation that intervenes at the right point
  • Experience mentoring junior engineers and contributing to team development
  • Experience applying AI/ML to operational challenges — intelligent alerting, automated diagnosis, predictive incident detection, or conversational operations interfaces
  • Familiarity with the AstraZeneca technology estate or regulated pharmaceutical environments
  • Experience operating platforms that serve AI/ML workloads (LLM inference, model serving, data pipelines)
  • ITIL, SRE, or operational excellence certifications or equivalent practical frameworks
  • Experience working within multi-tier support structures with clear escalation paths
  • Infrastructure-as-code expertise (Terraform, CloudFormation)
  • Awareness to GxP and audit trail awareness
  • Minimum of three days per week from the office

Desired Qualifications

  • Experience applying AI/ML to operational challenges — intelligent alerting, automated diagnosis, predictive incident detection, or conversational operations interfaces
  • Familiarity with the AstraZeneca technology estate or regulated pharmaceutical environments
  • Experience operating platforms that serve AI/ML workloads (LLM inference, model serving, data pipelines)
  • ITIL, SRE, or operational excellence certifications or equivalent practical frameworks
  • Experience working within multi-tier support structures with clear escalation paths
  • Infrastructure-as-code expertise (Terraform, CloudFormation)
  • Awareness to GxP and audit trail awareness

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce