Manulife

Lead Platform Reliability Engineer, Global AI Platform & Solutions

Manulife
Apply
1 hour ago
Toronto, CanadaStaff+

Responsibilities

  • Define SLOs and SLIs, track operational budgets, reduce MTTR, plan capacity, and tune autoscaling.
  • Build and maintain logging, metrics, tracing, alerting, dashboards, and operational runbooks.
  • Participate in on-call incident response, including triage, mitigation, root-cause analysis, postmortems, and corrective actions.
  • Develop self-service platform capabilities, AIOps/MLOps/GitOps/CI/CD pipelines, and automation for provisioning, upgrades, backups, and other operations.
  • Manage clusters, networks, storage, and policies using Terraform and Ansible while preventing configuration drift.
  • Enforce identity and access controls, secrets management, supply-chain security, regulatory controls, and AI/data governance requirements.
  • Optimize resource usage and cloud spend through rightsizing, autoscaling, reservations, spot capacity, and safe progressive delivery.
  • Treat the platform as a product by defining service catalogs, developer experience improvements, operational SLAs, and roadmap alignment.
  • Operate scalable backend services supporting high-traffic agent interactions, retrieval operations, and real-time execution flows.
  • Collaborate with global engineering, security, AI governance, risk, and audit teams to meet cross-geography and data-residency requirements.

Requirements

  • Bachelor's degree in Computer Science or Engineering, or equivalent experience and demonstrated skills.
  • 5–8 years of experience in DevOps, platform engineering, or production operations.
  • Proven experience operating large-scale distributed systems and participating in on-call support.
  • Experience with Azure, Kubernetes, containers, CI/CD, and observability stacks in cloud-native environments.
  • Programming experience with Python and/or Java, Scala, or TypeScript for backend services and automation.
  • Understanding of AI solutions, LLM systems, retrieval architectures, embeddings, vector stores, prompt and tool orchestration, and agent workflows.
  • Knowledge of API design, asynchronous workflows, concurrency, reliability engineering, SLOs, error budgets, and performance tuning.
  • Familiarity with authentication and authorization, data protection, audit logging, model governance, security, compliance, and governance for AI/data systems.
  • Ability to collaborate across global teams and translate business requirements into platform capabilities and operational SLAs.
  • Preferred qualifications include ITIL and ITSM certification, Azure Administrator or Azure DevOps certification, CKA or CKS certification, and HashiCorp Terraform Associate certification.

Benefits

  • Hybrid working arrangement in Toronto, Ontario.
  • Flexible environment supporting learning, career growth, well-being, and inclusion.
  • Eligible employees may receive health, dental, mental health, vision, disability, life, AD&D, adoption/surrogacy, wellness, and employee/family assistance benefits.
  • Retirement savings plans, including pension and a global share ownership plan with employer matching contributions, plus financial education and counseling resources.
  • Paid holidays, vacation, personal and sick days, and statutory leaves in Canada.
  • Eligibility for incentive programs and incentive compensation tied to business and individual performance.

Categories

DevOpsSite Reliability
Manulife

About Manulife

10,000+ employees
Contact me