Radiant

Infrastructure Tooling & Observability Engineer( UK)

Radiant
Apply
3 months ago
London, United KingdomSenior

Responsibilities

  • Design, build, and evolve internal tooling and observability platforms for large-scale distributed infrastructure operations.
  • Transform high-volume logs, metrics, and events into actionable insight and improve visibility, alerting quality, and operational decision-making.
  • Translate SRE reliability requirements into production-ready software for incident detection, prevention, and automated remediation.
  • Automate environment provisioning, cluster onboarding, inventory management, lifecycle workflows, and other infrastructure operations.
  • Build tooling for capacity management, performance testing, benchmarking, and automated performance-result collection and analysis.
  • Contribute to Continual Service Improvement initiatives by identifying operational inefficiencies and delivering durable engineering solutions.
  • Work with SRE, infrastructure engineering, and Platform Engineering teams to embed observability and reliability into core platform workflows.
  • Integrate and extend systems written in Ruby/Rails and Go.
  • Develop and maintain infrastructure automation workflows using Ansible and AWX.
  • Support CI/CD-driven operational tooling using GitHub Actions and self-hosted runners.

Requirements

  • A degree in Computer Science or Software Engineering, or equivalent experience.
  • 6–8 years of experience in infrastructure engineering, DevOps, SRE, and/or software engineering roles focused on operational systems.
  • Recent experience building or maintaining production infrastructure tooling or platform systems in a DevOps or software engineering role.
  • Experience working in large-scale or distributed infrastructure environments.
  • Strong programming ability in Ruby/Rails, Go, or similar systems languages, with the ability to work across multiple languages and codebases.
  • Hands-on experience with Ansible and AWX or similar infrastructure automation and orchestration tools.
  • Strong experience with the Grafana observability stack, including Prometheus, Loki, Mimir, and Grafana Alloy.
  • Familiarity with SNMP and syslog.
  • Production experience with Kubernetes or similar orchestration platforms.
  • Understanding of API design and integration patterns, particularly REST-based services and service-to-service communication.
  • Experience building and maintaining CI/CD pipelines with GitHub Actions and self-hosted runners.
  • Strong understanding of monitoring, alerting, capacity planning, incident response, and other operational reliability concepts.
  • Preferred qualifications include Kubernetes Certified Administrator certification, cloud-native observability training or industry conference participation, CompTIA+ Security Qualifications, and LPI/LPIC certification.

Benefits

  • Exposure to large-scale distributed infrastructure systems.
  • Opportunity to shape foundational internal platforms.
  • Collaborative, engineering-led culture with strong ownership.
  • High-impact work spanning observability, automation, and reliability.
  • Close partnership with SRE and infrastructure engineering teams.
  • Fast-moving environment where tooling directly improves operational performance.

Tech Stack

AnsibleGitHub ActionsGoGrafanaKubernetesPrometheusRubyRuby on Rails

Categories

Radiant

About Radiant

501-1,000 employees
Contact me