
Senior Site Reliability Engineer I
Electronic Arts2 hours ago
Hyderābād, IndiaSenior
Responsibilities
- Architect, implement, and manage resilient, scalable, highly available infrastructure systems.
- Automate operations, deployments, monitoring, and other manual processes to reduce toil and improve reliability.
- Create observability solutions and dashboards for proactive issue detection and remediation.
- Lead critical incident response, root cause analysis, corrective actions, runbooks, playbooks, and post-incident reviews.
- Oversee infrastructure performance, capacity planning, scalability, resource utilization, latency, and reliability across hybrid and cloud environments.
- Define and improve Service Level Objectives, Service Level Indicators, and monitoring strategies based on the Four Golden Signals.
- Provide technical leadership and mentorship to SRE teams and cross-functional engineering groups.
- Establish documentation, facilitate technical learning, and promote blameless postmortems and knowledge sharing.
- Develop SRE strategy, tooling roadmaps, automation frameworks, and continuous-improvement initiatives.
- Collaborate with security and compliance teams on secure configurations, vulnerability remediation, access controls, and DevSecOps practices.
- Engage executive and business stakeholders and represent SRE in governance forums, audits, and architecture reviews.
Requirements
- Bachelor’s or Master’s degree in Computer Science, Information Technology, or a related field.
- 12–15 years of total IT experience, including at least 8 years in SRE, DevOps, or large-scale systems engineering.
- Strong proficiency in Linux/Unix system administration and internals.
- Experience with AWS, Azure, or GCP cloud platforms.
- Advanced scripting and automation skills using Python, Go, PowerShell, or Bash.
- Hands-on experience with Docker, Kubernetes, and service meshes such as Istio.
- Expertise with monitoring and observability tools including Prometheus, Grafana, ELK, Datadog, Splunk, Zabbix, and Nagios.
- Expertise with configuration management and infrastructure-as-code tools including Ansible, Terraform, Chef, and Puppet.
- Strong knowledge of networking, load balancing, databases, and distributed systems.
- Experience with enterprise-scale incident response, problem management, capacity planning, reliability, redundancy, and disaster recovery.
- Excellent analytical, communication, leadership, mentoring, stakeholder management, and cross-functional collaboration skills.
- Preferred qualifications include experience establishing SRE frameworks or centers of excellence, REST API development and integration, database query optimization, governance and compliance frameworks, AIOps, self-healing systems, machine-learning-driven monitoring, organizational reliability transformations, and DevOps or SRE-related industry or open-source participation.
Benefits
- Benefits may include healthcare coverage, mental well-being support, retirement savings, paid time off, family leave, complimentary games, career support, and community wellness programs.
Tech Stack
AnsibleAWSAzureBashChefDatadogDockerGoGoogle Cloud PlatformGrafanaIstioKubernetesLinuxNagiosPowerShellPrometheusPuppetPythonSplunkTerraform
Categories
DevOpsSite Reliability
About Electronic Arts
Electronic Arts creates next-level entertainment experiences that inspire players and fans around the world. Here, everyone is part of the story. Part of a community that connects across the globe. A team where creativity thrives, new perspectives are invited, and ideas matter. Regardless of your role, team, or location, this is a place where everyone makes play happen. Join us.