18 days ago
Stoke-on-Trent, United KingdomSenior
Responsibilities
- Develop and maintain resilient tools, operational APIs, and automation for system management.
- Use orchestration and scripting to reduce manual activity, operational toil, and inconsistency.
- Write code, telemetry, and instrumentation that improve service reliability and observability.
- Build dashboards and operational views using Grafana, Splunk, New Relic, and related platforms.
- Configure and manage Cloudflare edge services using Infrastructure as Code and integrate edge telemetry with observability platforms.
- Diagnose incidents across edge, network, platform, application, dependency, and origin layers.
- Participate in live incident response, post-mortems, and root-cause analysis.
- Maintain monitoring, alerting, APM, analytics, and PagerDuty workflows.
- Drive initiatives that improve reliability, observability, performance, and continuous improvement across teams.
- Mentor colleagues, share knowledge, and collaborate with IT Operations on business-value-focused tooling.
Requirements
- Software engineering background with Python, Golang, JavaScript, or a similar language.
- Knowledge of modern development practices, including testing, source control, and delivery lifecycles.
- Understanding of SRE principles, including SLIs, SLOs, reliability measurement, and incident management.
- Hands-on experience with observability tools such as OpenTelemetry, Splunk, New Relic, Grafana, or PagerDuty.
- Proficiency in shell scripting for automation and system management.
- Experience with Infrastructure as Code, including Terraform and Ansible.
- Knowledge of Cloudflare or a comparable edge platform, including DNS, CDN, WAF, DDoS protection, and traffic management.
- Ability to troubleshoot distributed systems across edge, network, platform, application, dependency, and origin layers.
- Experience in a large-scale, 24/7 enterprise where uptime, performance, and stability are critical.
- Practical experience using LLM platforms and coding assistants safely to improve productivity, quality, and root-cause analysis.
Benefits
- Eligible for the company’s hybrid work-from-home policy.
- Inclusive workplace with opportunities for employee growth and development.
- Recruitment-process adjustments and accommodations are available upon request.
Tech Stack
Categories
Site Reliability
