Elsevier

Senior Site Reliability Engineer

Elsevier
Apply
1 day ago
London, United KingdomSenior

Responsibilities

  • Create monitoring queries, establish service-level baselines, and develop observability dashboards and configuration as code.
  • Support incident response, on-call operations, root-cause analyses, post-mortems, failovers, rollbacks, and operational runbooks.
  • Implement automation and execute code in production environments to reduce operational toil and improve reliability.
  • Participate in disaster recovery and non-production availability, reliability, and recoverability testing.
  • Support reliability architecture, infrastructure topology, deployment workflows, performance benchmarking, and production readiness reviews.
  • Operate and improve platforms, CI/CD processes, cloud infrastructure, and application migration workflows.
  • Support the deployment, monitoring, security, and reliability of services integrating AI tools.
  • Enable engineering teams through documentation, consultancy, troubleshooting, and secure self-service practices.

Requirements

  • Advanced Terraform expertise, including modules, providers, state management, lifecycle controls, drift detection, safe refactoring, remote state, locking, and cross-stack dependencies.
  • Hands-on experience managing production multi-account, multi-region AWS environments using ECS, RDS, ALB, VPC, IAM, Route 53, ECR, S3, Lambda, DynamoDB, SQS, Secrets Manager, KMS, and CloudWatch.
  • Experience building and troubleshooting reusable GitHub Actions workflows, OIDC authentication, approval gates, runners, Terraform deployments, application deployments, and migration pipelines.
  • Strong knowledge of Docker, ECS Fargate, ECS task definitions and services, IAM roles, health checks, autoscaling, ALB integration, and deployment rollbacks.
  • Proficiency in AWS networking, VPCs, ALBs, Route 53, ACM/TLS, IAM, OIDC, Secrets Manager, KMS, and cloud security best practices.
  • Expertise in incident response, observability, logs, metrics, alarms, deployment history, root-cause analysis, rollback decisions, and operational runbooks.
  • Strong Linux and Git fundamentals with Bash or Python scripting for AWS CLI automation, CI/CD, and operational tooling.
  • Hands-on experience integrating and operating AI services and APIs in production, including monitoring, reliability, and security practices.
  • Ability to support multiple engineering teams, troubleshoot across infrastructure and application layers, document solutions, and enable secure self-service practices.

Benefits

  • Flexible working hours.
  • Comprehensive pension plan.
  • Generous vacation entitlement and the option for sabbatical leave.
  • Maternity, paternity, adoption, and family care leave.
  • Internal communities and networks.
  • Various employee discounts.
  • Recruitment introduction reward.
  • Employee Assistance Program.
  • Wellbeing initiatives, shared parental leave, and study assistance.

Tech Stack

Amazon DynamoDBBashDockerGitGitHub ActionsLinuxPythonTerraform

Categories

Site Reliability
Elsevier

About Elsevier

10,000+ employees

Elsevier builds scientific and medical information products and analytics platforms for researchers, clinicians, and educators. Its portfolio includes journals and databases such as The Lancet, Cell, ScienceDirect, Scopus, and ClinicalKey, sold primarily via institutional subscriptions, licensing, and open-access publishing fees. Founded in 1880 and headquartered in Amsterdam, Elsevier operates as a business of RELX, serving universities, hospitals, corporations, and government research agencies worldwide.

Contact me