1 hour ago
Toronto, CanadaStaff+
Responsibilities
- Define SLOs and SLIs, track operational budgets, reduce MTTR, plan capacity, and tune autoscaling.
- Build and maintain logging, metrics, tracing, alerting, dashboards, and operational runbooks.
- Participate in on-call incident response, including triage, mitigation, root-cause analysis, postmortems, and corrective actions.
- Develop self-service platform capabilities, AIOps/MLOps/GitOps/CI/CD pipelines, and automation for provisioning, upgrades, backups, and other operations.
- Manage clusters, networks, storage, and policies using Terraform and Ansible while preventing configuration drift.
- Enforce identity and access controls, secrets management, supply-chain security, regulatory controls, and AI/data governance requirements.
- Optimize resource usage and cloud spend through rightsizing, autoscaling, reservations, spot capacity, and safe progressive delivery.
- Treat the platform as a product by defining service catalogs, developer experience improvements, operational SLAs, and roadmap alignment.
- Operate scalable backend services supporting high-traffic agent interactions, retrieval operations, and real-time execution flows.
- Collaborate with global engineering, security, AI governance, risk, and audit teams to meet cross-geography and data-residency requirements.
Requirements
- Bachelor's degree in Computer Science or Engineering, or equivalent experience and demonstrated skills.
- 5–8 years of experience in DevOps, platform engineering, or production operations.
- Proven experience operating large-scale distributed systems and participating in on-call support.
- Experience with Azure, Kubernetes, containers, CI/CD, and observability stacks in cloud-native environments.
- Programming experience with Python and/or Java, Scala, or TypeScript for backend services and automation.
- Understanding of AI solutions, LLM systems, retrieval architectures, embeddings, vector stores, prompt and tool orchestration, and agent workflows.
- Knowledge of API design, asynchronous workflows, concurrency, reliability engineering, SLOs, error budgets, and performance tuning.
- Familiarity with authentication and authorization, data protection, audit logging, model governance, security, compliance, and governance for AI/data systems.
- Ability to collaborate across global teams and translate business requirements into platform capabilities and operational SLAs.
- Preferred qualifications include ITIL and ITSM certification, Azure Administrator or Azure DevOps certification, CKA or CKS certification, and HashiCorp Terraform Associate certification.
Benefits
- Hybrid working arrangement in Toronto, Ontario.
- Flexible environment supporting learning, career growth, well-being, and inclusion.
- Eligible employees may receive health, dental, mental health, vision, disability, life, AD&D, adoption/surrogacy, wellness, and employee/family assistance benefits.
- Retirement savings plans, including pension and a global share ownership plan with employer matching contributions, plus financial education and counseling resources.
- Paid holidays, vacation, personal and sick days, and statutory leaves in Canada.
- Eligibility for incentive programs and incentive compensation tied to business and individual performance.
Tech Stack
Categories
DevOpsSite Reliability
