Location
toronto, on, Canada
Posted
July 15, 2026
Job Description
Join as an Expert AI Operations Engineer, ensuring the success of AI platforms through reliability and operational excellence. Oversee incident management and operational metrics in a compliant framework.
This pivotal role involves the administration of the organization's AI platform and its lifecycle. You will play a critical part in monitoring platform health, managing service readiness, and coordinating operations across various AI use cases. Your work will uphold the integration of AI technologies safely and effectively within the organization.
Key Responsibilities: • Ensure operational reliability across AI platforms • Monitor and report on service performance metrics • Facilitate incident response and service stability • Validate environment readiness for AI applications • Enhance automation and observability for efficiency
Requirements: • Degree in Computer Science, Engineering, or related discipline • 5–7 years in site reliability engineering or cloud...
This pivotal role involves the administration of the organization's AI platform and its lifecycle. You will play a critical part in monitoring platform health, managing service readiness, and coordinating operations across various AI use cases. Your work will uphold the integration of AI technologies safely and effectively within the organization.
Key Responsibilities: • Ensure operational reliability across AI platforms • Monitor and report on service performance metrics • Facilitate incident response and service stability • Validate environment readiness for AI applications • Enhance automation and observability for efficiency
Requirements: • Degree in Computer Science, Engineering, or related discipline • 5–7 years in site reliability engineering or cloud...