14 days ago
Bengaluru, IndiaSenior
Responsibilities
- Design, build, and operate scalable multi-cloud and hybrid infrastructure across AWS, GCP, and on-premises data centers.
- Own EKS and GKE Kubernetes platforms, including cluster lifecycle, multi-tenancy, networking, storage, multi-cluster patterns, and autoscaling.
- Implement infrastructure as code with Terraform and Pulumi and manage GitOps-based delivery using ArgoCD and Flux.
- Deploy progressive delivery patterns across game services and operate service mesh and Kubernetes networking capabilities.
- Build and operate observability using Prometheus, Grafana, Datadog, and OpenTelemetry.
- Define SLI, SLO, and error-budget policies and create effective alerting.
- Lead chaos engineering, incident response, and post-mortems focused on systemic improvements.
- Automate self-service provisioning, remediation, scaling, server configuration, and other repetitive operational work.
- Harden CI/CD pipelines and embed secrets management and policy-as-code at the platform layer.
- Promote SRE practices, conduct reliability reviews, create runbooks, author engineering RFCs, and influence architectural decisions across 2K studios.
Requirements
- 5+ years of experience in SRE, platform engineering, or equivalent production-scale infrastructure work.
- Deep Kubernetes experience in cloud environments, preferably with EKS or GKE, including networking, storage, and multi-cluster patterns.
- Strong infrastructure-as-code skills with Terraform and/or Pulumi, plus hands-on experience with Helm, Terragrunt, and GitOps tooling.
- Experience with AWS, GCP, VMware, and bare-metal servers.
- Experience configuring servers with Ansible, Puppet, and AWS Systems Manager.
- Experience with Datadog, Prometheus, Grafana, and OpenTelemetry.
- Fluency with SLI, SLO, and error-budget practices and their operationalization within engineering teams.
- Production-quality programming experience in Go, Python, or TypeScript for tools, automation, and internal libraries.
- Ability to debug Linux internals, TCP/IP networking, DNS, and TLS at the system level.
- Experience leading incident response and post-mortems with systemic follow-through.
- Preferred experience includes live-service games or large-scale consumer internet systems serving millions of concurrent users.
- Preferred experience includes Istio, Cilium, advanced Kubernetes networking, FinOps, cloud-scale resource management, AI and agentic development, cloud certifications, and mentoring SREs or leading reliability groups.
Benefits
- Company culture emphasizing creativity, innovation, efficiency, diversity, and philanthropy.
- Professional growth opportunities and a collaborative environment within a global entertainment company.
- Corporate boot camps, company parties, office gaming spaces, game release events, monthly socials, and team challenges.
- Discretionary bonus, provident fund contributions, medical insurance with top-up options, online doctor consultation access, employee assistance, life assurance, personal accident insurance, childcare services, and 20 days of holiday plus statutory holidays.
- Gym reimbursement up to INR1150 per month, wellbeing program, charitable giving program, learning platforms, employee discounts, and free games and events.
- Onsite role, as indicated by the #LI-Onsite designation.
Tech Stack
AnsibleAWSDatadogGitHub ActionsGoGoogle Cloud PlatformGrafanaHelmIstioJenkinsKubernetesLinuxPrometheusPuppetPythonTerraformTypeScript
Categories
DevOpsSite Reliability
