18 hours ago
London, United KingdomSenior
Responsibilities
- Create monitoring queries, service-level baselines, observability dashboards, SLOs, and error budgets.
- Participate in incident response, on-call support, post-mortems, root-cause analyses, failovers, and rollbacks.
- Design and implement automation, execute code in production, and reduce operational toil.
- Test availability, reliability, recoverability, and performance in non-production environments and support production readiness reviews.
- Support deployment, monitoring, reliability, and security for services integrating AI tools.
- Create infrastructure topology drawings, deployment workflows, operational-health reports, and reliability documentation.
- Support application migrations to standard platforms and provide direction and consultancy for platform and paved-road features.
- Analyze and improve SDLC and CI/CD processes while enabling secure self-service practices for multiple engineering teams.
Requirements
- Advanced Terraform expertise, including modules, providers, state management, lifecycle controls, drift detection, safe refactoring, remote state, locking, and cross-stack dependencies.
- Hands-on experience managing production, multi-account, multi-region AWS environments using ECS, RDS, ALB, VPC, IAM, Route 53, ECR, S3, Lambda, DynamoDB, SQS, Secrets Manager, KMS, and CloudWatch.
- Experience building and troubleshooting reusable GitHub Actions workflows, OIDC authentication, approval gates, runners, Terraform deployments, application deployments, and migration pipelines.
- Strong knowledge of Docker, ECS task definitions and services, IAM roles, health checks, autoscaling, ALB integration, and deployment rollbacks, including ECS Fargate.
- Proficiency in AWS networking and security, including VPCs, ALBs, Route 53, ACM/TLS, IAM, OIDC, Secrets Manager, KMS, and cloud security practices.
- Strong incident-response and observability skills using logs, metrics, alarms, deployment history, root-cause analysis, rollback decisions, and operational runbooks.
- Strong Linux and Git fundamentals with Bash and Python scripting for AWS CLI automation, CI/CD, and operational tooling.
- Hands-on experience integrating and operating AI services and APIs in production, including monitoring, reliability, and security practices.
- Ability to support multiple engineering teams, troubleshoot across infrastructure and application layers, document solutions, and enable secure self-service practices.
Benefits
- Flexible working hours.
- Comprehensive pension plan.
- Generous vacation entitlement and the option for sabbatical leave.
- Maternity, paternity, adoption, and family care leave.
- Internal communities and networks.
- Various employee discounts.
- Recruitment introduction reward.
- Global Employee Assistance Program.
- Wellbeing initiatives, shared parental leave, and study assistance.
Categories
Site Reliability
About RELX
RELX provides information-based analytics, research content, and decision tools for scientists, legal and risk professionals, insurers, and governments. Through segments including Elsevier (STM), LexisNexis (legal and risk), and RX (exhibitions), it sells subscriptions, data services, and software platforms used in 180+ countries. Headquartered in London, it is a public company listed in London and Amsterdam.