2K

Senior Site Reliability Engineer

2K
Apply
14 days ago
Bengaluru, IndiaSenior

Responsibilities

  • Design, build, and operate scalable multi-cloud and hybrid infrastructure across AWS, GCP, and on-premises data centers.
  • Own EKS and GKE Kubernetes platforms, including cluster lifecycle, multi-tenancy, networking, storage, multi-cluster patterns, and autoscaling.
  • Implement infrastructure as code with Terraform and Pulumi and manage GitOps-based delivery using ArgoCD and Flux.
  • Deploy progressive delivery patterns across game services and operate service mesh and Kubernetes networking capabilities.
  • Build and operate observability using Prometheus, Grafana, Datadog, and OpenTelemetry.
  • Define SLI, SLO, and error-budget policies and create effective alerting.
  • Lead chaos engineering, incident response, and post-mortems focused on systemic improvements.
  • Automate self-service provisioning, remediation, scaling, server configuration, and other repetitive operational work.
  • Harden CI/CD pipelines and embed secrets management and policy-as-code at the platform layer.
  • Promote SRE practices, conduct reliability reviews, create runbooks, author engineering RFCs, and influence architectural decisions across 2K studios.

Requirements

  • 5+ years of experience in SRE, platform engineering, or equivalent production-scale infrastructure work.
  • Deep Kubernetes experience in cloud environments, preferably with EKS or GKE, including networking, storage, and multi-cluster patterns.
  • Strong infrastructure-as-code skills with Terraform and/or Pulumi, plus hands-on experience with Helm, Terragrunt, and GitOps tooling.
  • Experience with AWS, GCP, VMware, and bare-metal servers.
  • Experience configuring servers with Ansible, Puppet, and AWS Systems Manager.
  • Experience with Datadog, Prometheus, Grafana, and OpenTelemetry.
  • Fluency with SLI, SLO, and error-budget practices and their operationalization within engineering teams.
  • Production-quality programming experience in Go, Python, or TypeScript for tools, automation, and internal libraries.
  • Ability to debug Linux internals, TCP/IP networking, DNS, and TLS at the system level.
  • Experience leading incident response and post-mortems with systemic follow-through.
  • Preferred experience includes live-service games or large-scale consumer internet systems serving millions of concurrent users.
  • Preferred experience includes Istio, Cilium, advanced Kubernetes networking, FinOps, cloud-scale resource management, AI and agentic development, cloud certifications, and mentoring SREs or leading reliability groups.

Benefits

  • Company culture emphasizing creativity, innovation, efficiency, diversity, and philanthropy.
  • Professional growth opportunities and a collaborative environment within a global entertainment company.
  • Corporate boot camps, company parties, office gaming spaces, game release events, monthly socials, and team challenges.
  • Discretionary bonus, provident fund contributions, medical insurance with top-up options, online doctor consultation access, employee assistance, life assurance, personal accident insurance, childcare services, and 20 days of holiday plus statutory holidays.
  • Gym reimbursement up to INR1150 per month, wellbeing program, charitable giving program, learning platforms, employee discounts, and free games and events.
  • Onsite role, as indicated by the #LI-Onsite designation.

Tech Stack

AnsibleAWSDatadogGitHub ActionsGoGoogle Cloud PlatformGrafanaHelmIstioJenkinsKubernetesLinuxPrometheusPuppetPythonTerraformTypeScript

Categories

DevOpsSite Reliability
2K

About 2K

1,001-5,000 employees
Contact me