2 hours ago
London, United KingdomMid Level
Responsibilities
- Extend and maintain OpenTelemetry SDKs, libraries, and collectors.
- Build and support instrumentation paths across Kubernetes and non-Kubernetes workloads.
- Embed observability standards and golden paths across platform and application teams.
- Use AI-assisted engineering to improve productivity and engineering value.
- Scale and improve telemetry backend systems and pipelines.
- Improve incident response through better telemetry coverage.
- Participate in the out-of-hours on-call rotation.
Requirements
- Strong problem-solving skills and the ability to turn standards into usable libraries, examples, and guidance.
- Programming skills in C# and/or Python with a good understanding of software architecture.
- Hands-on experience with OpenTelemetry SDKs, instrumentation, and collectors.
- Proficiency in Kubernetes and cloud-native observability.
- Understanding of metrics, logs, and traces and how engineering teams use them to debug and operate services.
- Experience with DevOps tooling such as Terraform, ArgoCD, Helm, or Jenkins.
- Interest in AI engineering and SRE practices for improving incident response and root-cause analysis.
- Desirable experience includes Prometheus/PromQL, VictoriaMetrics, OpenSearch, Grafana, auto-instrumentation, distributed tracing, structured logging, trace/metric correlation, Kafka, or telemetry pipeline architectures.
Benefits
- Highly competitive compensation plus an annual discretionary bonus.
- Lunch provided through Just Eat for Business and access to a dedicated barista bar.
- 35 days of annual leave.
- 9% company pension contributions.
- Informal dress code and excellent work/life balance.
- Comprehensive healthcare and life assurance.
- Cycle-to-work scheme.
- Monthly company events.
- London-based role at the company's London headquarters with participation in an out-of-hours on-call rotation.
