London Stock Exchange Group

Senior Engineer - Site Reliability Engineering

London Stock Exchange Group
Apply
1 month ago
Raleigh, NC, USASenior

Responsibilities

  • Establish SRE foundations for new projects by building environments, monitoring, alerting, and operational-readiness practices.
  • Define and implement observability standards across metrics, logs, traces, and SLIs/SLOs.
  • Design and evolve monitoring and alerting solutions using tools such as Datadog, Prometheus, Grafana, ELK, and OpenTelemetry.
  • Drive reliability improvements through incident reduction, performance tuning, resilient design patterns, and reduced operational toil.
  • Partner with Security teams to meet compliance, security, risk-management, and cloud-cost objectives.
  • Lead project-to-BAU SRE handovers through documentation, readiness validation, and strong operational practices.
  • Provide technical leadership and mentorship while shaping engineering standards and supporting a culture of learning.
  • Support major incidents and critical issues when required.

Requirements

  • Bachelor’s degree in Computer Science or a related field.
  • At least 5 years of hands-on technical experience in SRE, Platform Engineering, Infrastructure, or related roles.
  • Strong experience with Azure, including AKS, Azure Container Apps, Virtual Machines, VNet, Entra ID, and Azure managed services.
  • Hands-on experience with Kubernetes and containerized platforms.
  • Proven experience designing and operating observability platforms for monitoring, logging, and alerting.
  • Hands-on experience with Datadog for metrics, logs, APM, and alerting.
  • Strong understanding of SRE principles, including SLOs, error budgets, incident management, and reliability engineering.
  • Understanding of cloud security principles and experience collaborating with security teams.
  • Experience with cloud-cost optimization strategies and tooling.
  • Preferred experience with AWS, multi-cloud or hybrid environments, Infrastructure as Code such as Terraform or CloudFormation, and large-scale, complex, or regulated environments.
  • Preferred knowledge of vector databases, RAG architectures, Generative AI, LLM platforms such as Claude and Amazon Bedrock, and integrating AI with observability stacks.
  • Strong technical authority, collaboration, communication, problem-solving, and incident-response skills.

Benefits

  • Healthcare, retirement planning, paid volunteering days, and wellbeing initiatives are offered.
  • The role is based in London and is part of LSEG’s global, collaborative organization.
  • LSEG supports charitable involvement through the LSEG Foundation and employee fundraising and volunteering.
  • The company provides reasonable accommodation for religious practices and mental or physical disabilities.

Tech Stack

AWSAzureDatadogGrafanaKubernetesPrometheusTerraform

Categories

Site Reliability
London Stock Exchange Group

About London Stock Exchange Group

10,000+ employees
Contact me