3 months ago
London, United KingdomSenior
Responsibilities
- Design, build, and evolve internal tooling and observability platforms for large-scale distributed infrastructure operations.
- Transform high-volume logs, metrics, and events into actionable insight and improve visibility, alerting quality, and operational decision-making.
- Translate SRE reliability requirements into production-ready software for incident detection, prevention, and automated remediation.
- Automate environment provisioning, cluster onboarding, inventory management, lifecycle workflows, and other infrastructure operations.
- Build tooling for capacity management, performance testing, benchmarking, and automated performance-result collection and analysis.
- Contribute to Continual Service Improvement initiatives by identifying operational inefficiencies and delivering durable engineering solutions.
- Work with SRE, infrastructure engineering, and Platform Engineering teams to embed observability and reliability into core platform workflows.
- Integrate and extend systems written in Ruby/Rails and Go.
- Develop and maintain infrastructure automation workflows using Ansible and AWX.
- Support CI/CD-driven operational tooling using GitHub Actions and self-hosted runners.
Requirements
- A degree in Computer Science or Software Engineering, or equivalent experience.
- 6–8 years of experience in infrastructure engineering, DevOps, SRE, and/or software engineering roles focused on operational systems.
- Recent experience building or maintaining production infrastructure tooling or platform systems in a DevOps or software engineering role.
- Experience working in large-scale or distributed infrastructure environments.
- Strong programming ability in Ruby/Rails, Go, or similar systems languages, with the ability to work across multiple languages and codebases.
- Hands-on experience with Ansible and AWX or similar infrastructure automation and orchestration tools.
- Strong experience with the Grafana observability stack, including Prometheus, Loki, Mimir, and Grafana Alloy.
- Familiarity with SNMP and syslog.
- Production experience with Kubernetes or similar orchestration platforms.
- Understanding of API design and integration patterns, particularly REST-based services and service-to-service communication.
- Experience building and maintaining CI/CD pipelines with GitHub Actions and self-hosted runners.
- Strong understanding of monitoring, alerting, capacity planning, incident response, and other operational reliability concepts.
- Preferred qualifications include Kubernetes Certified Administrator certification, cloud-native observability training or industry conference participation, CompTIA+ Security Qualifications, and LPI/LPIC certification.
Benefits
- Exposure to large-scale distributed infrastructure systems.
- Opportunity to shape foundational internal platforms.
- Collaborative, engineering-led culture with strong ownership.
- High-impact work spanning observability, automation, and reliability.
- Close partnership with SRE and infrastructure engineering teams.
- Fast-moving environment where tooling directly improves operational performance.
