2 hours ago
Montréal, CanadaSenior
Responsibilities
- Lead high-severity incident investigations, root-cause analysis, post-mortems, and durable reliability improvements.
- Implement end-to-end traces, metrics, logs, anomaly detection, synthetics, contract tests, and topology-aware health models.
- Build policy-driven auto-healing and reliability tooling, including circuit breakers, throttling, retries, progressive delivery, safe rollbacks, chaos, disaster recovery, and runbooks.
- Define and govern user-centric SLIs, SLOs, SLAs, error budgets, alert strategies, capacity models, and reliability reporting.
- Coach application support, infrastructure support, and incident management teams and standardize playbooks and training.
- Operate and improve reliability across Azure, AWS, GCP, on-premises environments, Kubernetes, service meshes, data systems, and global networking.
- Apply AI to causal detection, anomaly detection, reliability copilots, and monitoring AI systems for reliability and cost.
Requirements
- 8+ years of experience in SRE, platform, infrastructure, or software engineering operating large-scale production systems across multi-cloud and on-premises environments.
- Strong proficiency with OpenTelemetry, Dynatrace, Elastic/ELK, reliability engineering practices, progressive delivery, Kubernetes, service meshes, and platform resilience.
- Experience with data and event systems, replication, snapshots/PITR, CDC, Kafka, RabbitMQ, Pub/Sub, dead-letter queues, and reprocessing.
- Knowledge of DNS, load balancers, CDN/edge, TLS/mTLS, BGP, and global traffic management.
- Strong software engineering skills in at least one of Go, Python, or TypeScript.
- Experience with Terraform, Argo CD or Flux, GitOps, policy as code, chaos engineering, game days, and disaster recovery exercises.
- Excellent written, visual, and verbal communication skills, including coaching teams and presenting to technical and business audiences.
- Bilingual French and English, with the ability to regularly interact with English-speaking clients and colleagues across Canada.
- Must be eligible to work in Canada; Canadian work experience is not required.
Benefits
- Flexible work arrangements and a hybrid work model.
- Possibility to purchase up to five additional days off per year.
- Physical and mental wellbeing benefits, including telemedicine and a Wellness account.
- Share plan and other savings opportunities, including company matching through the Employee Share Purchase Plan.
- Defined benefit pension plan with an opportunity for guaranteed income for life.
- Annual bonus plan with a target based on base salary.
- The role is based on a 35-hour workweek.
Tech Stack
Apache KafkaArgo CDAWSAzureGoGoogle Cloud PlatformIstioKibanaKubernetesPythonRabbitMQTerraformTypeScript
Categories
DevOpsSite Reliability
