2 months ago
Montréal, CanadaMid Level
Responsibilities
- Own application-level infrastructure for the Monetization domain, including Helm charts, Terraform configurations, Kubernetes deployments, runtime configuration, networking, and integrations.
- Design and maintain SLOs, SLIs, monitors, alerts, and dashboards using Datadog and OpenTelemetry-based tooling.
- Build and evolve CI/CD pipelines with deployment and rollback automation.
- Perform capacity planning, load testing, performance tuning, and performance regression investigation for launches, sales events, and regional rollouts.
- Lead production readiness reviews and define production-readiness standards for new services and major changes.
- Support incident response, deep investigations, post-mortems, runbooks, and follow-up reliability improvements.
- Develop automation and tooling to reduce operational toil, including runbook automation, deployment helpers, and operational scripts.
- Maintain the domain reliability roadmap and participate in product planning, refinements, architecture reviews, and company-wide SRE standards.
- Participate in the SRE duty rotation and contribute improvements to shared SRE-operated systems.
Requirements
- At least 3 years of SRE, DevOps, or platform engineering experience involving on-call or incident response, SLO and monitoring ownership, deployment pipelines, and production infrastructure.
- Backend software development experience, including shipping services and writing production-quality automation in a language such as Go or PHP.
- Hands-on Kubernetes experience with Helm, manifests, deployment strategies, managed Kubernetes, and application performance and networking troubleshooting.
- Experience building observability systems with monitors, dashboards, and SLOs/SLIs; Datadog is preferred, while Prometheus, Grafana, and OpenTelemetry experience is relevant.
- Infrastructure-as-code experience with Terraform or Terragrunt.
- Experience with GCP, including IAM, networking, and managed services.
- Experience building and maintaining CI/CD pipelines with GitLab CI and/or GitHub Actions.
- Programming or scripting proficiency in Python, Go, Bash, or comparable tools for automation and tooling.
- Practical incident response and post-mortem experience with a record of driving reliability improvements.
- Strong collaboration and communication skills for daily work with product development teams.
- Experience in payments, fintech, e-commerce, or gaming and other high-traffic transactional systems is valued.
- Kubernetes, Google Cloud Platform, or HashiCorp certifications are nice to have.
Benefits
- Hybrid embedded role within the SRE organization.
- Medical, dental, and vision coverage.
- Paid time off.
- Personalized career roadmap.
- Training and educational opportunities for professional development.
- Support for employees' physical, mental, and emotional well-being and their families.
Tech Stack
BashDatadogGitHub ActionsGitLab CI/CDGoGoogle Cloud PlatformGrafanaHelmKubernetesPHPPrometheusPythonTerraform
Categories
Site Reliability
