5 months ago
Toronto, Canada or Montréal, CanadaStaff+
Responsibilities
- Own the enterprise resiliency reference architecture and define availability, latency, durability, RTO, and RPO requirements.
- Establish architecture governance through design reviews, production gates, policy-as-code, scorecards, and automated controls.
- Standardize blue/green deployments, traffic shifting, health gates, progressive cutovers, rollback, and zero-downtime migrations.
- Lead enterprise chaos engineering, failure injection, game days, disaster recovery drills, and automated failover practices.
- Define production readiness standards, runbooks, dependency maps, failover topologies, and service-level objectives.
- Drive observability and SRE practices including OpenTelemetry, distributed tracing, error budgets, and reliability dashboards.
- Architect disaster recovery and cyber-resilience strategies including immutable backups, recovery validation, and ransomware-resistant segmentation.
- Guide platform and data resiliency across Kubernetes, service meshes, replication, geo-distribution, and event streaming.
- Enable reliable AI/GenAI systems and AI-driven operations through monitoring, guardrails, anomaly detection, predictive modeling, and remediation.
- Serve as the principal resilience authority by mentoring teams, leading councils and forums, and communicating tradeoffs to executives and engineers.
Requirements
- 10+ years of experience in SRE, platform, infrastructure, or systems architecture with large-scale production-critical environments.
- Proven experience across Azure, AWS, Google Cloud Platform, and on-premises environments.
- Experience with multi-region traffic management, global load balancing, DNS/BGP, TLS/mTLS, and CDN/edge patterns.
- Experience with Kubernetes ecosystems, service meshes, autoscaling, readiness and liveness controls, and topology constraints.
- Knowledge of observability stacks, distributed tracing, correlation, and topology modeling.
- Experience with consensus and replication, partitioning, point-in-time recovery, snapshots, change data capture, caches, and databases.
- Experience with infrastructure as code, automation, GitOps, policy-as-code, CI/CD, blue/green deployments, canary releases, and progressive delivery.
- Experience leading chaos engineering, disaster recovery orchestration, and automated enterprise-scale failover.
- AI/GenAI experience including model serving, vector stores, retrieval pipelines, guardrails, model monitoring, feature pipelines, lineage, and AI-assisted operations.
- Strong software engineering skills in Go, Python, or TypeScript and strong systems thinking.
- Excellent written, visual, and verbal communication skills with executive presence.
- Candidates located in Quebec must be bilingual because the role regularly interacts with English-speaking colleagues across Canada.
- Candidates must be eligible to work in Canada; no Canadian work experience is required.
Benefits
- Flexible work arrangements and a hybrid work model.
- Possibility to purchase up to five additional days off per year.
- Physical and mental wellbeing benefits including telemedicine and a wellness account.
- Share plan and other savings opportunities, including an employee share purchase plan with 50% matching of net shares.
- Defined benefit pension plan and the opportunity for guaranteed income for life.
- Annual bonus plan with a 15% target and potential payout of up to double the target.
- Accessible and inclusive workplace with accommodation support and equal opportunity practices.
Tech Stack
Argo CDAWSAzureGoGoogle Cloud PlatformGrafanaIstioKubernetesPrometheusPythonRedisTerraformTypeScript
Categories
DevOpsSite Reliability
