Site Reliability Developer 1
$85,000–$115,000 year
RemoteToronto, Ontario, Canada or Canada
Job Summary
Support key ITIL processes including incident, request, problem, and change management while defining runbooks and standard operating procedures. Field operational requests from the Application Support team and internal stakeholders to triage and solve issues within defined SLAs, ensuring an excellent customer experience. Maintain live services by measuring availability, latency, and system health using DataDog and CloudWatch, identifying root causes, and championing fixes across the organization. Work with infrastructure-as-code via Terraform and Ansible in an agile environment, collaborating on features and reporting on performance metrics. Participate in an on-call rotation covering 10am–10pm every 2–3 weeks.
Required Qualifications
- Bachelor's degree in computer science, Software engineering or equivalent experience
- 2+ years of experience in an IT Operational, DevOps, SRE, or Software Engineering role
- Experience with cloud computing (AWS and Azure) services
- Developing-level of knowledge with the management and setup of cloud infrastructure
- You can write code - in any language
- You have implemented your work in a production environment and can back it up with examples
- Experience with tools and platforms such as: Ansible, Build/Release Pipelines, Docker, Github, Terraform etc
- Developing-level of knowledge with distributed systems in the cloud using observability and telemetry for oversight of code deployments and service level objectives (SLOs)
- Developing experience with the operational aspects of software systems using telemetry, centralized logging, and alerting with tools such as: CloudWatch, Datadog, Prometheus, etc
- Participating in an on-call rotation every 2–3 weeks, covering 10am–10pm from Monday through Sunday
Desired Qualifications
- Experience with Terraform
- Experience with Ansible
- Experience with Jenkins
- Experience with RDS MySQL
- Experience with Redshift
- Experience with Redshift Spectrum
- Experience with MongoDB
- Experience with Elasticsearch
- Experience with Kinesis
- Experience with SQS
- Experience with RabbitMQ
- Experience with Python
- Experience with Java
- Experience with Dropwizard
- Experience with Spring Boot
- Experience with Hibernate
- Experience with TypeScript
- Experience with JavaScript
- Experience with React
- Experience with Context Api
- Experience with Hooks
- Experience with Redux
- Experience with DataDog
- Experience with CloudWatch
- Experience with Prometheus
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.