EQ Bank logo
EQ BankPosted 1 month ago

Senior AI Platform Operations Engineer

On-siteToronto, Ontario, Canada

Part TimeSenior LevelLarge

Job Summary

Administer and operate the AI platform to ensure availability, performance, and resilience across environments, integrations, and supporting infrastructure. Monitor platform health using dashboards, logs, metrics, and alerts, leading operational triage, escalation coordination, and post-incident reviews to strengthen service stability. Enable approved AI use cases into production by validating dependencies, completing operational readiness checklists, and executing structured service transitions. Implement observability capabilities including telemetry, logging, and traces while analyzing operational data to identify anomalies and root-cause patterns. Execute governance controls for AI solutions, ensuring usage and access compliance, auditability, and alignment with enterprise security policies. Maintain operational visibility of AI assets and support ongoing audits to ensure accurate cost attribution and lifecycle status.

Required Qualifications

  • University degree in Computer Science, Engineering, Information Technology, or a related field, or equivalent practical experience
  • 5-7 years of experience in platform operations, site reliability engineering, DevOps, cloud operations, or enterprise IT operations
  • Strong experience supporting production platforms and services, including monitoring, incident response, problem management, service restoration, and operational reporting
  • Experience with cloud platforms, observability, automation, configuration management, and integration patterns, including Azure Automation runbooks (PowerShell/Python), Azure AI, Copilot integrations, AKS, virtual networks (hub-and-spoke), and App Service
  • Expertise with observability tools such as Azure Monitor, Application Insights, and Grafana
  • Experience with CI/CD and automation tools such as Azure DevOps, GitHub Actions, and Logic Apps
  • Knowledge of configuration management and infrastructure-as-code tools such as Bicep, Terraform, Azure Policy, Key Vault, and relevant open-source technologies
  • Knowledge of integration and event-driven technologies such as API Management, open-source API tools, Service Bus, Event Grid, and Apache Kafka
  • Working knowledge of platform-supporting data and search services such as Elastic, Azure AI Search, and Cosmos DB
  • Knowledge of security, privacy, audit, and compliance considerations relevant to enterprise AI and platform operations

Desired Qualifications

  • Knowledge of enterprise network, edge security, and related internal platforms such as DNA, Fortinet, and Akamai is an asset
  • Working knowledge of AI/ML operational concepts, including model lifecycle support, telemetry, governance controls, human-in-the-loop practices, and production monitoring
  • Strong understanding of ITIL/ITSM processes, including change, release, incident, problem, configuration, and service reporting practices
  • Analytical and structured thinker with strong troubleshooting, root-cause analysis, prioritization, and continuous improvement skills
  • Strong service orientation, professional maturity, and the ability to collaborate effectively across operations, engineering, security, risk, data, and business teams
  • Experience creating technical documentation, operational procedures, support playbooks, dashboards, and user guidance materials

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce