Xsolla

Site Reliability Engineer (Monetization)

Xsolla
Apply
2 months ago
Montréal, CanadaMid Level

Responsibilities

  • Own application-level infrastructure for the Monetization domain, including Helm charts, Terraform configurations, Kubernetes deployments, runtime configuration, networking, and integrations.
  • Design and maintain SLOs, SLIs, monitors, alerts, and dashboards using Datadog and OpenTelemetry-based tooling.
  • Build and evolve CI/CD pipelines with deployment and rollback automation.
  • Perform capacity planning, load testing, performance tuning, and performance regression investigation for launches, sales events, and regional rollouts.
  • Lead production readiness reviews and define production-readiness standards for new services and major changes.
  • Support incident response, deep investigations, post-mortems, runbooks, and follow-up reliability improvements.
  • Develop automation and tooling to reduce operational toil, including runbook automation, deployment helpers, and operational scripts.
  • Maintain the domain reliability roadmap and participate in product planning, refinements, architecture reviews, and company-wide SRE standards.
  • Participate in the SRE duty rotation and contribute improvements to shared SRE-operated systems.

Requirements

  • At least 3 years of SRE, DevOps, or platform engineering experience involving on-call or incident response, SLO and monitoring ownership, deployment pipelines, and production infrastructure.
  • Backend software development experience, including shipping services and writing production-quality automation in a language such as Go or PHP.
  • Hands-on Kubernetes experience with Helm, manifests, deployment strategies, managed Kubernetes, and application performance and networking troubleshooting.
  • Experience building observability systems with monitors, dashboards, and SLOs/SLIs; Datadog is preferred, while Prometheus, Grafana, and OpenTelemetry experience is relevant.
  • Infrastructure-as-code experience with Terraform or Terragrunt.
  • Experience with GCP, including IAM, networking, and managed services.
  • Experience building and maintaining CI/CD pipelines with GitLab CI and/or GitHub Actions.
  • Programming or scripting proficiency in Python, Go, Bash, or comparable tools for automation and tooling.
  • Practical incident response and post-mortem experience with a record of driving reliability improvements.
  • Strong collaboration and communication skills for daily work with product development teams.
  • Experience in payments, fintech, e-commerce, or gaming and other high-traffic transactional systems is valued.
  • Kubernetes, Google Cloud Platform, or HashiCorp certifications are nice to have.

Benefits

  • Hybrid embedded role within the SRE organization.
  • Medical, dental, and vision coverage.
  • Paid time off.
  • Personalized career roadmap.
  • Training and educational opportunities for professional development.
  • Support for employees' physical, mental, and emotional well-being and their families.

Tech Stack

BashDatadogGitHub ActionsGitLab CI/CDGoGoogle Cloud PlatformGrafanaHelmKubernetesPHPPrometheusPythonTerraform

Categories

Site Reliability
Xsolla

About Xsolla

1,001-5,000 employees
Contact me