Tech Talent International logo
Tech Talent InternationalPosted 1 month ago

Senior Site Reliability Engineer (SRE) – Automation & Observability

$110,000–$120,000 year

On-siteMontréal, Quebec, Canada

Full TimeSenior LevelSmall

Job Summary

Monitor, manage, and proactively improve the availability and performance of production environments across application and infrastructure layers. Plan and implement application deployments, load testing, and configuration changes while collaborating with development teams to ensure operational readiness. Contribute to architecture reviews, solution design, and migrations, and identify opportunities to enhance operational efficiency through automation initiatives. Respond to incidents and resolve issues quickly, often under pressure, with participation in on-call rotations and after-hours support. Provide constructive feedback on performance and capacity, assist in documenting architectures, and maintain focus on target solutions within a global team.

Required Qualifications

  • 5–7 years of experience in a similar role
  • Experience providing multidisciplinary technical support within a team environment
  • Practical knowledge of performance and capacity management across: Applications, Databases, Networks
  • Strong automation skills and mindset
  • Strong Linux/Unix administration skills
  • Good knowledge of Windows environments
  • Strong knowledge of Docker and Kubernetes
  • Understanding of cloud-based platforms and solutions
  • Good understanding of enterprise infrastructure, firewalls, and networking concepts
  • Knowledge of load-balancing technologies
  • Strong understanding of networking fundamentals
  • Experience with APIs
  • Familiarity with CyberArk or HashiCorp Vault
  • Experience with SQL Server
  • Experience with Oracle
  • Exposure to NoSQL databases
  • Experience configuring application monitoring tools such as Dynatrace
  • Experience with: Jenkins, Bitbucket, Artifactory, Ansible, ArgoCD
  • Knowledge of software development and scripting methodologies
  • Demonstrated programming ability in languages such as Python
  • Good understanding of ITIL processes
  • Understanding of user and server authentication mechanisms that enable automated deployment cycles while maintaining strong security controls
  • Strong problem-solving abilities
  • Team-oriented mindset
  • Customer-focused approach
  • Participation in on-call rotations and after-hours support
  • Roles starts off 5 days in office for 1st 3 months, then turns into hybrid setup 3 days onsite, 2 days from home

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce