McKesson logo
McKessonPosted 1 week ago

Senior Lead Engineer, Platform Operations & Observability

$94,400–$157,300 year

HybridMontréal, Quebec, Canada

Full TimeSenior LevelBachelors DegreeLarge

Job Summary

Lead monitoring and observability strategies across enterprise applications, platforms, and services for the Canada B2C Digital Solution. Design, implement, and optimize dashboards, alerts, telemetry, logging, and performance monitoring solutions. Drive incident management processes, major incident response, escalation coordination, service restoration activities, and postmortem investigations. Conduct root cause analysis (RCA) investigations and lead corrective and preventive action planning. Partner with engineering teams to improve platform reliability, resiliency, scalability, and operational readiness. Lead change management reviews and promote safe deployment and release practices. Provide technical leadership, coaching, and mentoring to engineers while establishing engineering best practices and influencing architecture, automation, and CI/CD initiatives.

Required Qualifications

  • 7+ years of professional experience in Software Engineering, Site Reliability Engineering, Platform Engineering, DevOps, or related technical roles
  • Bachelor's degree in Computer Science, Engineering, Information Technology, or equivalent experience
  • Experience supporting large-scale production environments and enterprise applications
  • Hands-on experience with monitoring, observability, logging, alerting, and application performance monitoring tools
  • Proven experience leading incident management and production support activities
  • Experience performing root cause analysis and implementing preventive solutions
  • Experience with CI/CD, automation, DevOps practices, and software delivery pipelines
  • Experience with microservices, APIs, distributed systems, and cloud-based architectures
  • Lead production readiness reviews and operational acceptance activities prior to major releases
  • Hybrid, two mandatory days at the Dobrin Office (Usually on Monday and Wednesday)
  • Ability to work at a computer for extended periods and participate in virtual collaboration activities
  • Participation in on-call support and prod deployment rotations out of business hours may be required based on organizational needs

Desired Qualifications

  • Experience with tools such as Dynatrace, Prometheus, Dotcom Monitor or similar observability platforms
  • Experience with Kubernetes, containers, and cloud platforms such as Azure
  • Knowledge of ITIL-aligned incidents, problems, and change management practices
  • Experience defining SLAs, MTTR and service reliability metrics
  • Demonstrated technical leadership and mentorship of engineering teams
  • Experience operating in regulated or highly compliant environments
  • Experience driving platform modernization and operational excellence initiatives
  • Strong analytical, troubleshooting, continuous improvement and stakeholder communication skills
  • Fair understanding and mastery of Ai tools (Copilot, Rovo) and Ai agents

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce