
Engineer - Site Reliability Engineering
London Stock Exchange Group2 hours ago
St. Louis, MO, USASenior
Responsibilities
- Maintain Service Level Objectives for owned systems and improve availability, latency, and system health.
- Write automation to scale systems sustainably, prevent service issues, and restore service quickly.
- Partner with development teams to improve reliability, observability, and release velocity.
- Participate in on-call rotations, incident response, postmortems, and root cause analysis and resolution.
- Advocate for engineering practices that support scalable, reliable, and performant services.
- Support cloud migration through architectural reviews, operational acceptance testing, and configuration of Datadog dashboards and metrics.
Requirements
- Bachelor’s degree in computer science or a related technical field involving software or systems engineering, or equivalent practical experience.
- Experience with object-oriented programming languages such as Java, C#, Python, or Go.
- Experience with Unix/Linux and Windows operating systems.
- Hands-on experience with Azure, AWS, or GCP.
- Preferred: 8–10 years of industry experience.
- Preferred: experience with DevOps concepts, algorithms and data structures, observability practices, infrastructure as code, identity and access management, and application security.
- Experience with Datadog, BigPanda, Terraform, or EntraID is relevant to the team’s observability, cloud infrastructure, and IAM environment.
Benefits
- Healthcare benefits, retirement planning, paid volunteering days, and wellbeing initiatives.
- Collaborative, inclusive culture with opportunities for continuous learning and development.
- Equal-opportunity employer with reasonable accommodations for religious practices, mental health needs, and physical disabilities.
Categories
Site Reliability