
Technical Lead - Site Reliability Engineering
London Stock Exchange Group1 month ago
Raleigh, NC, USAStaff+
Responsibilities
- Establish SRE foundations for new projects, including environments, monitoring, alerting, and operational readiness.
- Define and implement observability standards covering metrics, logs, traces, and SLIs/SLOs.
- Design monitoring and alerting solutions and drive reliability improvements through incident reduction, performance tuning, and resilient patterns.
- Collaborate with Architecture, Engineering, Security, and Platform teams on reliable, scalable, secure system design and delivery.
- Support project-to-BAU SRE handovers through documentation, readiness checks, and operational practices.
- Drive cloud cost optimization and efficiency initiatives using data-driven analysis.
- Mentor engineers, shape engineering standards, and promote continuous learning.
- Support major incidents and critical issues when required.
Requirements
- Bachelor’s degree in Computer Science or a related field.
- 10+ years of hands-on technical experience in SRE, Platform Engineering, Infrastructure, or related roles.
- Strong experience with Azure, including AKS, Azure Container Apps, Virtual Machines, VNet, and Azure managed services.
- Hands-on experience with Kubernetes and containerized platforms and a strong background in Linux system administration.
- Experience designing and operating observability platforms for monitoring, logging, and alerting.
- Hands-on Datadog experience covering metrics, logs, APM, and alerting.
- Strong understanding of SRE principles, including SLOs, error budgets, incident management, and reliability engineering.
- Experience collaborating with architecture and engineering teams on system design and delivery.
- Understanding of cloud security principles and experience working with security teams.
- Experience with cloud cost optimization strategies and tooling.
- Experience integrating AI with observability stacks such as Prometheus, Grafana, ELK, and OpenTelemetry for proactive issue detection.
- Preferred qualifications include AWS experience, multi-cloud or hybrid environment support, Infrastructure as Code with Terraform or CloudFormation, experience in large-scale or regulated environments, and knowledge of vector databases, RAG architectures, generative AI, and LLM platforms such as Claude and Amazon Bedrock.
- Strong technical authority, communication, collaboration, problem-solving, and incident leadership skills.
Benefits
- Healthcare, retirement planning, paid volunteering days, and wellbeing initiatives.
- Collaborative and creative culture with opportunities to contribute ideas and professional development.
- Equal opportunity employer with reasonable accommodations for religious practices, mental health needs, and physical disabilities.
- Global working environment across LSEG’s 25,000 colleagues in 65 countries.