
Senior Engineer - Site Reliability Engineering
London Stock Exchange Group1 month ago
Raleigh, NC, USASenior
Responsibilities
- Establish SRE foundations for new projects by building environments, monitoring, alerting, and operational-readiness practices.
- Define and implement observability standards across metrics, logs, traces, and SLIs/SLOs.
- Design and evolve monitoring and alerting solutions using tools such as Datadog, Prometheus, Grafana, ELK, and OpenTelemetry.
- Drive reliability improvements through incident reduction, performance tuning, resilient design patterns, and reduced operational toil.
- Partner with Security teams to meet compliance, security, risk-management, and cloud-cost objectives.
- Lead project-to-BAU SRE handovers through documentation, readiness validation, and strong operational practices.
- Provide technical leadership and mentorship while shaping engineering standards and supporting a culture of learning.
- Support major incidents and critical issues when required.
Requirements
- Bachelor’s degree in Computer Science or a related field.
- At least 5 years of hands-on technical experience in SRE, Platform Engineering, Infrastructure, or related roles.
- Strong experience with Azure, including AKS, Azure Container Apps, Virtual Machines, VNet, Entra ID, and Azure managed services.
- Hands-on experience with Kubernetes and containerized platforms.
- Proven experience designing and operating observability platforms for monitoring, logging, and alerting.
- Hands-on experience with Datadog for metrics, logs, APM, and alerting.
- Strong understanding of SRE principles, including SLOs, error budgets, incident management, and reliability engineering.
- Understanding of cloud security principles and experience collaborating with security teams.
- Experience with cloud-cost optimization strategies and tooling.
- Preferred experience with AWS, multi-cloud or hybrid environments, Infrastructure as Code such as Terraform or CloudFormation, and large-scale, complex, or regulated environments.
- Preferred knowledge of vector databases, RAG architectures, Generative AI, LLM platforms such as Claude and Amazon Bedrock, and integrating AI with observability stacks.
- Strong technical authority, collaboration, communication, problem-solving, and incident-response skills.
Benefits
- Healthcare, retirement planning, paid volunteering days, and wellbeing initiatives are offered.
- The role is based in London and is part of LSEG’s global, collaborative organization.
- LSEG supports charitable involvement through the LSEG Foundation and employee fundraising and volunteering.
- The company provides reasonable accommodation for religious practices and mental or physical disabilities.
Tech Stack
Categories
Site Reliability