15 hours ago
Dublin, IrelandSenior
Responsibilities
- Design, implement, and maintain highly available Azure infrastructure, including Azure Kubernetes Service, Virtual Machines, Functions, App Services, Storage, Networking, Key Vault, and Azure Monitor.
- Define and maintain Service Level Indicators, Service Level Objectives, and Error Budgets to improve reliability and availability.
- Lead root cause analysis, post-incident reviews, resilience engineering, disaster recovery, multi-region failover, and high-availability initiatives.
- Design Terraform Infrastructure as Code components and automate operational processes using Python, PowerShell, Go, or Bash.
- Operate and optimize AKS clusters, establish container security standards, implement ArgoCD-based GitOps practices, and manage upgrades, capacity, and workloads.
- Develop monitoring, logging, alerting, distributed tracing, and application performance monitoring strategies using Grafana and Azure Monitor.
- Build and support auditable CI/CD pipelines using GitHub Actions and Azure DevOps, including blue/green and canary deployments.
- Provide technical leadership, mentor SRE I and SRE II engineers, promote SRE practices, and participate in on-call rotations and major incident management.
Requirements
- Experience in Site Reliability Engineering, Cloud Engineering, DevOps, or Platform Engineering supporting large-scale production environments.
- Expert knowledge of Microsoft Azure and strong Kubernetes experience, preferably with Azure Kubernetes Service.
- Advanced Terraform skills and experience designing reusable Infrastructure as Code components.
- Strong Git and GitOps practices, with experience using GitHub Actions and Azure DevOps.
- Strong Linux administration, scripting, and programming skills using Python, PowerShell, Go, or Bash.
- Experience with monitoring and observability platforms, distributed tracing, application performance monitoring, and proactive alerting.
- Skills in incident management, root cause analysis, capacity planning, performance optimization, disaster recovery testing, and production operations.
- Technical leadership, strategic thinking, problem solving, communication, collaboration, customer focus, and an automation-first approach.
- Azure, Kubernetes, or Terraform certifications are welcomed.
Benefits
- Dublin, Ireland location at Rockfield Central.
- Country-specific benefits are offered.
- Participation in an on-call rotation is required.
- The employer provides an accessible hiring process and disability or other accommodation support.
Categories
DevOpsSite Reliability
About RELX
RELX provides information-based analytics, research content, and decision tools for scientists, legal and risk professionals, insurers, and governments. Through segments including Elsevier (STM), LexisNexis (legal and risk), and RX (exhibitions), it sells subscriptions, data services, and software platforms used in 180+ countries. Headquartered in London, it is a public company listed in London and Amsterdam.