14 days ago
Responsibilities
- Design, build, and operate scalable multi-cloud and hybrid infrastructure across AWS, GCP, and on-premises data centers.
- Own Kubernetes platforms end to end, including cluster lifecycle, multi-tenancy, networking, storage, and autoscaling.
- Implement infrastructure as code and GitOps workflows using Terraform, Pulumi, ArgoCD, and Flux.
- Build and operate observability, alerting, SLI/SLO, error-budget, and reliability systems.
- Lead chaos engineering, incident response, and post-mortems focused on systemic fixes.
- Automate provisioning, remediation, scaling, CI/CD, secrets management, and policy enforcement.
- Promote SRE practices across studios through reliability reviews, runbooks, embedded collaboration, and mentoring.
- Shape architecture and author engineering RFCs that advance the platform.
Requirements
- 5+ years of experience in SRE, platform engineering, or equivalent production-scale infrastructure work.
- Deep Kubernetes experience in cloud environments, preferably EKS or GKE, including networking, storage, and multi-cluster patterns.
- Strong Terraform and/or Pulumi experience, plus hands-on experience with Helm, Terragrunt, and GitOps tooling.
- Experience with AWS, GCP, VMware, bare-metal servers, Ansible, Puppet, and AWS Systems Manager.
- Experience with Datadog, Prometheus, Grafana, and OpenTelemetry observability stacks.
- Fluency with SLI/SLO and error-budget practices, including operationalizing them within engineering teams.
- Production-quality coding experience in Go, Python, or TypeScript.
- Strong Linux internals, TCP/IP networking, DNS, and TLS debugging skills.
- Incident response and post-mortem leadership experience with systemic follow-through.
- Preferred: live-service game or large-scale consumer internet experience, service mesh and advanced Kubernetes networking expertise, FinOps experience, AI and agentic development experience, relevant cloud or Kubernetes certifications, and experience mentoring SREs or leading reliability groups.
Benefits
- Base compensation is provided with additional potential bonus, equity awards, and medical, financial, and other benefits for regular employees.
- Hybrid work arrangement in Burnaby, Canada.
- Candidates must be legally authorized to work in Canada without current or future employer sponsorship.
- Reasonable accommodation is available for applicants and employees with disabilities.
Tech Stack
AnsibleAWSDatadogGitHub ActionsGoGoogle Cloud PlatformGrafanaHelmIstioJenkinsKubernetesLinuxPrometheusPuppetPythonTerraformTypeScript
Categories
DevOpsSite Reliability
H1B sponsorship
No H1B sponsorship
All candidates must be legally authorized to work in Canada without requiring current or future employer sponsorship
2K’s past filings provide context. They do not guarantee sponsorship for a current opening.
Verified yearly history is not available for this company.
Job fields receiving H1B certifications
A job-field breakdown is not available for this company.
Source: US Department of Labor