3 months ago
Remote, United KingdomStaff+
Responsibilities
- Own production reliability for high-SLA and complex customer environments.
- Define and evolve per-tenant SLOs and reliability models while reducing SLO budget burn.
- Design and implement automation, self-healing, auto-scaling, and other solutions to eliminate toil and improve reliability.
- Serve as a primary escalation point and participate in on-call and customer-impacting incident response.
- Lead incident investigations, post-incident reviews, and customer communications through bridge calls when necessary.
- Improve monitoring, alert quality, customer observability, and noisy escalation handling.
- Design fault-tolerant and scalable systems across the service lifecycle.
- Partner with Mimir, Loki, and Tempo product engineering squads on feature design, technical designs, roadmaps, and product strategy.
- Contribute to design documents, code reviews, and pull request reviews.
- Teach Site Reliability Engineering practices and mentor engineers.
Requirements
- At least 8 years of engineering experience, including at least 4 years in SRE, customer reliability engineering, or production engineering.
- Strong Kubernetes experience in AWS, GCP, or Azure and familiarity with Helm, Terraform, or Jsonnet.
- Experience operating multi-tenant production systems and designing and implementing SLOs.
- Experience with one or more programming languages such as Go, Python, or Java.
- Experience with Linux operating system internals and knowledge of networking, cloud storage, and scaling.
- Strong technical leadership experience, including leading projects, mentoring engineers, and serving as a force multiplier.
- Experience participating in incident response, conducting investigations, following up on actions, and writing high-quality post-incident reviews.
- Ability to reason about performance, scaling, failure modes, and production reliability.
- Ability to partner deeply with product engineering teams and work autonomously.
- Formal customer reliability engineering experience is strongly preferred.
Benefits
- 100% remote work opportunity for candidates in the UK, Sweden, Spain, or Germany.
- In-person onboarding is provided.
- 30 days of annual leave per year, including 3 Grafana Shutdown Days, subject to local legislation.
- Equity and bonus may be available.
- Career growth pathways, transparent communication, and an autonomy-focused global culture.
Categories
DevOpsSite Reliability
About Grafana
Grafana Labs builds open-source observability tools and a managed SaaS, Grafana Cloud, used by engineering and SRE teams to monitor, visualize, and analyze telemetry. Its products include Grafana dashboards plus Loki (logs), Tempo (traces), and Mimir (metrics), offered as cloud subscriptions and enterprise software. Privately held and headquartered in New York, it serves customers such as Anthropic, Bloomberg, NVIDIA, Microsoft, and Salesforce.
