
Site Reliability Engineer
London Stock Exchange Group2 months ago
St. Louis, MO, USASenior
Responsibilities
- Maintain SLOs for cloud-hosted systems while improving availability, latency, scalability, and system health.
- Design and implement automation for AWS and Azure environments to prevent service issues and enable rapid recovery.
- Partner with development teams to improve reliability, observability, and release velocity using cloud-native tools.
- Participate in on-call rotations, incident response, postmortems, and root cause analysis.
- Support scalable and reliable service deployment and operations across AWS and Azure.
- Enable cloud migration through architectural reviews, operational readiness testing, and observability dashboard configuration.
- Drive continuous learning and development in multi-cloud technologies.
Requirements
- Bachelor’s degree or equivalent experience in Computer Science or a related technical field.
- Proficiency in Shell and Python scripting.
- Hands-on Infrastructure as Code experience with Terraform, AWS CloudFormation, or Azure ARM templates.
- Experience with AWS services including EC2, S3, RDS, IAM, VPC, and Lambda, plus familiarity with Azure equivalents.
- Knowledge of Docker and Kubernetes, preferably EKS or AKS.
- Familiarity with Datadog, AWS CloudWatch, or Azure Monitor.
- Experience implementing and maintaining CI/CD pipelines using AWS CodePipeline, Azure DevOps, or similar tools.
- Experience with incident response, blameless postmortems, and Git workflows.
- Preferred: 2–5 years of proven experience.
- Preferred: understanding of the AWS Well-Architected Framework and Azure Architecture Center.
- Preferred: experience with cloud security guidelines and services such as IAM, KMS, and Secrets Manager.
- Preferred: familiarity with AWS RDS, DynamoDB, and Azure SQL, plus cloud-native architectures and DevOps concepts.
Benefits
- Healthcare benefits, retirement planning, paid volunteering days, and wellbeing initiatives.
- Collaborative and inclusive culture with opportunities for learning and continuous improvement.
- Global organization with sustainability initiatives and charitable volunteering opportunities.
Categories
DevOpsSite Reliability