London Stock Exchange Group

Technical Lead - Site Reliability Engineering

London Stock Exchange Group
Apply
1 month ago
Raleigh, NC, USAStaff+

Responsibilities

  • Establish SRE foundations for new projects, including environments, monitoring, alerting, and operational readiness.
  • Define and implement observability standards covering metrics, logs, traces, and SLIs/SLOs.
  • Design monitoring and alerting solutions and drive reliability improvements through incident reduction, performance tuning, and resilient patterns.
  • Collaborate with Architecture, Engineering, Security, and Platform teams on reliable, scalable, secure system design and delivery.
  • Support project-to-BAU SRE handovers through documentation, readiness checks, and operational practices.
  • Drive cloud cost optimization and efficiency initiatives using data-driven analysis.
  • Mentor engineers, shape engineering standards, and promote continuous learning.
  • Support major incidents and critical issues when required.

Requirements

  • Bachelor’s degree in Computer Science or a related field.
  • 10+ years of hands-on technical experience in SRE, Platform Engineering, Infrastructure, or related roles.
  • Strong experience with Azure, including AKS, Azure Container Apps, Virtual Machines, VNet, and Azure managed services.
  • Hands-on experience with Kubernetes and containerized platforms and a strong background in Linux system administration.
  • Experience designing and operating observability platforms for monitoring, logging, and alerting.
  • Hands-on Datadog experience covering metrics, logs, APM, and alerting.
  • Strong understanding of SRE principles, including SLOs, error budgets, incident management, and reliability engineering.
  • Experience collaborating with architecture and engineering teams on system design and delivery.
  • Understanding of cloud security principles and experience working with security teams.
  • Experience with cloud cost optimization strategies and tooling.
  • Experience integrating AI with observability stacks such as Prometheus, Grafana, ELK, and OpenTelemetry for proactive issue detection.
  • Preferred qualifications include AWS experience, multi-cloud or hybrid environment support, Infrastructure as Code with Terraform or CloudFormation, experience in large-scale or regulated environments, and knowledge of vector databases, RAG architectures, generative AI, and LLM platforms such as Claude and Amazon Bedrock.
  • Strong technical authority, communication, collaboration, problem-solving, and incident leadership skills.

Benefits

  • Healthcare, retirement planning, paid volunteering days, and wellbeing initiatives.
  • Collaborative and creative culture with opportunities to contribute ideas and professional development.
  • Equal opportunity employer with reasonable accommodations for religious practices, mental health needs, and physical disabilities.
  • Global working environment across LSEG’s 25,000 colleagues in 65 countries.

Tech Stack

AWSAzureDatadogGrafanaKubernetesLinuxPrometheusTerraform

Categories

DevOpsSite Reliability
London Stock Exchange Group

About London Stock Exchange Group

10,000+ employees
Contact me