
Senior Site Reliability Engineer, Data & Analytics
Blizzard Entertainment3 days ago
Base Salary
$103k - $190k/yr
Responsibilities
- Participate in on-call rotations, drive incidents to resolution, and lead blameless postmortems.
- Improve the reliability, scalability, performance, and cost of distributed data, analytics, ML training, and inference workloads.
- Support GPU-based ML training pipelines and inference services.
- Define and operate data and ML services on Kubernetes.
- Design and build automation, diagnostic tooling, runbooks, internal tools, and paved paths to reduce toil.
- Build and evolve centralized platform services, shared tooling, data integrations, and access-control patterns.
- Maintain infrastructure using Terraform and infrastructure-as-code practices.
- Improve CI/CD and GitOps workflows using Jenkins, GitHub Actions, and ArgoCD.
- Define and measure reliability with SLIs, SLOs, and error budgets.
- Run load tests, capacity modeling, and production validation.
- Partner with data, ML, platform, and other cross-functional engineering teams.
Requirements
- Experience operating reliable distributed systems in SRE, platform, or similar roles.
- Experience with data, analytics, ML, or large-scale distributed workloads.
- Strong knowledge of Linux, containers, Kubernetes, and cloud infrastructure.
- Experience building automation or internal tools with Python, Go, shell, or similar technologies.
- Experience with infrastructure-as-code such as Terraform.
- Experience with CI/CD or GitOps systems such as Jenkins, GitHub Actions, or ArgoCD.
- Familiarity with observability, including metrics, logs, traces, alerting, and incident response.
- Understanding of SRE concepts including SLIs, SLOs, error budgets, and postmortems.
- Experience improving reliability and efficiency through modern development and automation practices.
- Experience building internal tooling, automation, or developer productivity systems.
- Strong communication skills with technical and cross-functional partners.
- Preferred experience with ML training pipelines, model serving, GPU workloads, distributed systems, Kafka, Pub/Sub, Kubernetes-based environments, Prometheus, Grafana, GCP, or AWS.
Benefits
- Medical, dental, vision, health savings or reimbursement accounts, healthcare and dependent care spending accounts, life and disability insurance.
- 401(k) with Company match, tuition reimbursement, and charitable donation matching.
- Paid holidays and vacation, paid sick time, floating holidays, compassion and bereavement leave, and parental leave.
- Mental health and wellbeing programs, fitness programs, free and discounted games, and additional voluntary benefits.
- Relocation assistance may be available when the Company requires geographic relocation.
- The role is available in Irvine, California or Albany, New York with hybrid or on-site work, or fully remotely.
- Benefits are subject to eligibility requirements and may vary for part-time, temporary full-time employees, and interns.
Tech Stack
Apache KafkaAWSGitHub ActionsGoGoogle Cloud PlatformGrafanaJenkinsKubernetesLinuxPrometheusPythonTerraform
Categories
DevOpsSite Reliability
About Blizzard Entertainment
Blizzard Entertainment develops and publishes video games for PC and consoles, including World of Warcraft, Overwatch, Diablo, StarCraft, and Hearthstone, supported by Battle.net services and esports. Founded in 1991 (as Silicon & Synapse) and headquartered in Irvine, California, it operates within Activision Blizzard, which was acquired by Microsoft in 2023. Its business model spans premium games, expansions, a World of Warcraft subscription, and in-game purchases for live-service titles.