
Software Engineer III, Site Reliability Engineering - Ventures
Electronic Arts18 hours ago
Guildford, United KingdomSenior
Responsibilities
- Design and improve logging, metrics, tracing, dashboards, alerting, service-level indicators, and alert thresholds across Ventures environments.
- Lead incident response practices including triage, escalation, on-call rotations, runbooks, post-incident reviews, and follow-up actions.
- Establish cloud-spend visibility, cost-efficiency improvements, tagging standards, governance controls, and forecasting across AWS and GCP.
- Review and improve secrets management, IAM, least-privilege access, network configuration, and infrastructure security posture.
- Build and refine agentic workflows that inspect incidents, correlate telemetry, propose remediations, and support human-reviewed operations.
- Support Terraform infrastructure-as-code modules, existing CI/CD pipelines, and internally provided EA infrastructure and platform services.
- Document operational procedures and runbooks to enable independent team ownership after the contract ends.
Requirements
- 5+ years of professional SRE, DevOps, platform, or infrastructure engineering experience.
- Strong experience with observability tools such as Datadog, Grafana, Prometheus, OpenTelemetry, or CloudWatch.
- Production incident-response experience including on-call operations and post-incident reviews.
- Experience managing cloud costs at organizational scale, including tagging and governance practices.
- Understanding of cloud security fundamentals including secrets management, least-privilege access, and network segmentation.
- Strong hands-on AWS experience and working knowledge of GCP or demonstrated multi-cloud experience.
- Practical Terraform experience and familiarity with CI/CD tools such as GitHub Actions, GitLab CI, or Jenkins.
- Experience with Docker, Kubernetes, and scripting in Python, Bash, Go, or similar.
- Experience collaborating with internal platform or shared infrastructure teams in a larger organization.
- Hands-on experience using AI coding or operations agents and sound judgment about their appropriate use.
- Ability to become productive quickly in an existing environment and communicate clearly with technical and non-technical partners.
Benefits
- Six-month temporary full-time contract.
- Benefits may include healthcare coverage, mental well-being support, retirement savings, paid time off, family leave, complimentary games, and career and community wellness support, subject to local offerings.
Tech Stack
AWSBashDatadogDockerGitHub ActionsGitLab CI/CDGoGoogle Cloud PlatformGrafanaJenkinsKubernetesPrometheusPythonTerraform
Categories
Site Reliability
About Electronic Arts
Electronic Arts creates next-level entertainment experiences that inspire players and fans around the world. Here, everyone is part of the story. Part of a community that connects across the globe. A team where creativity thrives, new perspectives are invited, and ideas matter. Regardless of your role, team, or location, this is a place where everyone makes play happen. Join us.