2K

Senior Site Reliability Engineer

2K
Apply
8 hours ago

Responsibilities

  • Design, build, and operate scalable hybrid infrastructure across AWS and VMware vSphere using Terraform and Puppet.
  • Own EC2 fleets and VMware clusters, including capacity planning, patching, image pipelines, and autoscaling.
  • Architect AWS networking and operate F5 BIG-IP load balancers, firewalls, and global traffic management.
  • Operate MySQL and PostgreSQL production environments, including replication, failover, backup, recovery, upgrades, migrations, and performance tuning.
  • Build and run observability using Prometheus, Grafana, Datadog, and OpenTelemetry, and define SLI, SLO, and error-budget practices.
  • Lead chaos engineering, failover exercises, incident response, post-mortems, and systemic reliability improvements.
  • Automate provisioning, remediation, scaling, configuration management, F5 infrastructure, secrets management, policy enforcement, and CI/CD pipelines.
  • Promote SRE practices, contribute to architectural decisions, write engineering RFCs, and collaborate with network, database, architecture, and game development teams.

Requirements

  • 5+ years of experience in SRE, Platform Engineering, or equivalent production-scale infrastructure work.
  • Deep AWS experience with EC2, VPC networking, IAM, RDS/Aurora, S3, Route 53, ELB, and CloudWatch.
  • Strong VMware vSphere experience with ESXi, vCenter, vSAN and/or NSX, plus bare-metal server operations.
  • Hands-on F5 BIG-IP experience covering LTM, AFM, DNS/GTM, iRules, SSL/TLS offload, and high-availability pairs.
  • Production experience running MySQL and PostgreSQL, including replication, high availability, failover, backup and restore, upgrades, and performance tuning.
  • Strong Terraform and Packer experience, plus deep Puppet experience and hands-on Ansible and AWS Systems Manager experience.
  • Experience with Datadog, Prometheus, Grafana, and OpenTelemetry.
  • Practical fluency with SLIs, SLOs, and error budgets, including implementation within engineering teams.
  • Production-quality coding experience in Go, Python, or TypeScript for tools, automation, and internal libraries.
  • System-level debugging knowledge of Linux internals, TCP/IP networking, DNS, and TLS.
  • Incident response and post-mortem leadership with demonstrated systemic follow-through.
  • Preferred qualifications include live-service game or large-scale consumer internet experience, database reliability at scale, F5 automation, VMware-to-AWS migrations, FinOps, AI and agentic development, relevant AWS/VMware/F5/Puppet certifications, and experience mentoring SREs or leading reliability groups.

Benefits

  • Hybrid work arrangement, as indicated by the #LI-Hybrid designation.
  • 2K is an equal opportunity employer and provides reasonable accommodation for qualified individuals with disabilities.
  • The position does not provide visa sponsorship or assistance; candidates must already be legally authorized to work in the United States without current or future employer sponsorship.

Tech Stack

AnsibleAWSDatadogGitHub ActionsGoGrafanaJenkinsLinuxMySQLPostgreSQLPrometheusPuppetPythonTerraformTypeScript

Categories

Site Reliability
2K

About 2K

1,001-5,000 employees

2K develops and publishes video games for console, PC, and mobile players, selling premium titles and live-service content across major franchises. Its portfolio includes NBA 2K, BioShock, Borderlands, Sid Meier’s Civilization, XCOM, WWE 2K, and PGA TOUR 2K, built by in-house and partner studios. Founded in 2005 and headquartered in Novato, California, 2K is a publishing label wholly owned by Take-Two Interactive (NASDAQ: TTWO).

Contact me