Principal DevOps Engineer
On-siteToronto, Ontario, Canada
Job Summary
Drive initiatives to implement and enforce best practices for data streaming, processing, analytics, and monitoring infrastructure. Deploy and manage services on Kubernetes-based platforms such as Amazon EKS and Google Kubernetes Engine, while provisioning cloud infrastructure using Terraform to ensure security, scalability, and cost-efficiency. Maintain and optimize CI/CD pipelines using Jenkins, ArgoCD, and GitHub Enterprise Actions, and develop automation scripts in Python to support DevOps processes. Monitor system performance, troubleshoot issues, and implement SRE practices to enhance reliability and efficiency. Analyze cloud costs and implement strategies to reduce expenses while ensuring compliance with security policies. Collaborate with cross-functional teams to improve development workflows and influence data teams to align with best DevOps and SRE practices.
Required Qualifications
- 7+ years of experience in a DevOps, Site Reliability Engineering, or Cloud Infrastructure role
- Strong experience with AWS and GCP data services, including Kinesis, Glue, Pub/Sub, and Dataflow
- Proficiency in deploying and managing workloads on Kubernetes (EKS/GKE) in production environments
- Hands-on experience with Infrastructure-as-Code (IaC) using Terraform
- Expertise in CI/CD pipeline management using Jenkins, ArgoCD, and GitHub Enterprise Actions
- Programming skills in Python for automation and scripting
- Experience with observability and monitoring tools (e.g., Prometheus, Grafana, Datadog, or CloudWatch)
- Strong understanding of SRE principles, including performance monitoring, incident response, and reliability engineering
- Experience with cost optimization strategies for cloud infrastructure
- Self-motivated and driven, with a strong ability to influence and drive changes across multiple teams
- Ability to work collaboratively in an agile environment and support multiple teams
Desired Qualifications
- Experience with data lake architectures and big data processing frameworks (e.g., Apache Spark, Flink, Snowflake, BigQuery)
- Familiarity with event-driven architectures and message queues (e.g., Kafka, RabbitMQ)
- Experience with workflow orchestration tools such as Apache Airflow and Google Cloud Composer
- Knowledge of service mesh technologies like Istio
- Experience with GitOps workflows and Kubernetes-native tooling
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.