Electronic Arts

Senior Site Reliability Engineer I

Electronic Arts
Apply
2 hours ago
Hyderābād, IndiaSenior

Responsibilities

  • Architect, implement, and manage resilient, scalable, highly available infrastructure systems.
  • Automate operations, deployments, monitoring, and other manual processes to reduce toil and improve reliability.
  • Create observability solutions and dashboards for proactive issue detection and remediation.
  • Lead critical incident response, root cause analysis, corrective actions, runbooks, playbooks, and post-incident reviews.
  • Oversee infrastructure performance, capacity planning, scalability, resource utilization, latency, and reliability across hybrid and cloud environments.
  • Define and improve Service Level Objectives, Service Level Indicators, and monitoring strategies based on the Four Golden Signals.
  • Provide technical leadership and mentorship to SRE teams and cross-functional engineering groups.
  • Establish documentation, facilitate technical learning, and promote blameless postmortems and knowledge sharing.
  • Develop SRE strategy, tooling roadmaps, automation frameworks, and continuous-improvement initiatives.
  • Collaborate with security and compliance teams on secure configurations, vulnerability remediation, access controls, and DevSecOps practices.
  • Engage executive and business stakeholders and represent SRE in governance forums, audits, and architecture reviews.

Requirements

  • Bachelor’s or Master’s degree in Computer Science, Information Technology, or a related field.
  • 12–15 years of total IT experience, including at least 8 years in SRE, DevOps, or large-scale systems engineering.
  • Strong proficiency in Linux/Unix system administration and internals.
  • Experience with AWS, Azure, or GCP cloud platforms.
  • Advanced scripting and automation skills using Python, Go, PowerShell, or Bash.
  • Hands-on experience with Docker, Kubernetes, and service meshes such as Istio.
  • Expertise with monitoring and observability tools including Prometheus, Grafana, ELK, Datadog, Splunk, Zabbix, and Nagios.
  • Expertise with configuration management and infrastructure-as-code tools including Ansible, Terraform, Chef, and Puppet.
  • Strong knowledge of networking, load balancing, databases, and distributed systems.
  • Experience with enterprise-scale incident response, problem management, capacity planning, reliability, redundancy, and disaster recovery.
  • Excellent analytical, communication, leadership, mentoring, stakeholder management, and cross-functional collaboration skills.
  • Preferred qualifications include experience establishing SRE frameworks or centers of excellence, REST API development and integration, database query optimization, governance and compliance frameworks, AIOps, self-healing systems, machine-learning-driven monitoring, organizational reliability transformations, and DevOps or SRE-related industry or open-source participation.

Benefits

  • Benefits may include healthcare coverage, mental well-being support, retirement savings, paid time off, family leave, complimentary games, career support, and community wellness programs.

Tech Stack

AnsibleAWSAzureBashChefDatadogDockerGoGoogle Cloud PlatformGrafanaIstioKubernetesLinuxNagiosPowerShellPrometheusPuppetPythonSplunkTerraform

Categories

DevOpsSite Reliability
Electronic Arts

About Electronic Arts

10,000+ employees

Electronic Arts creates next-level entertainment experiences that inspire players and fans around the world. Here, everyone is part of the story. Part of a community that connects across the globe. A team where creativity thrives, new perspectives are invited, and ideas matter. Regardless of your role, team, or location, this is a place where everyone makes play happen. Join us.

Contact me