Responsibilities
- Own production reliability for high-SLA and complex customer environments.
- Define and evolve per-tenant SLOs and reliability models while reducing SLO budget burn.
- Design and implement automation, self-healing, auto-scaling, and other solutions that reduce operational toil.
- Improve customer observability, alert quality, scalability, and fault-tolerant design.
- Serve as an escalation point and participate in on-call, incident investigation, resolution, post-incident reviews, and customer bridge calls.
- Lead or contribute to design documents, code reviews, product strategy, roadmaps, and technical designs.
- Mentor engineers and teach Site Reliability Engineering practices that can be applied early in the development lifecycle.
- Partner with Mimir, Loki, and Tempo product engineering squads in an embedded model.
Requirements
- 8+ years of engineering experience, including 4+ years in SRE, customer reliability engineering, or production engineering.
- Strong Kubernetes experience in AWS, GCP, or Azure and familiarity with Helm, Terraform, or Jsonnet.
- Experience operating multi-tenant production systems and designing and implementing SLOs.
- Experience with one or more programming languages such as Go, Python, or Java.
- Experience with Linux operating system internals, networking, cloud storage, and scaling.
- Strong technical leadership, project leadership, mentoring, and cross-team partnership skills.
- Experience with incident response, post-incident reviews, troubleshooting, performance analysis, scaling, and failure-mode reasoning.
- Ability to work autonomously, communicate transparently, and collaborate deeply with product engineering teams.
Benefits
- 100% remote work arrangement for candidates in the UK, Sweden, Spain, or Germany.
- In-person onboarding.
- 30 days of annual leave per year, including 3 Grafana Shutdown Days, subject to local legislation.
- Equity and bonus eligibility, if applicable.
- Career growth pathways and a global, collaborative culture.
Tech Stack
Categories
About Grafana
Grafana Labs, the company behind the open observability cloud, is founded on the principles of open source, open standards, open ecosystems, and open culture. Grafana Cloud, our fully managed observability platform, is flexible and built for scale. With Grafana Cloud's actually useful AI, organizations can see, understand, and act on all their disparate data to move at the speed of their ambitions, while getting the visibility they need to run AI systems reliably and at scale. Today, more than 35 million users and 7,000+ customers – including Anthropic, Bloomberg, NVIDIA, Microsoft, and Salesforce – trust Grafana Labs to ensure reliability of their applications and systems, resolve incidents quickly, and optimize their telemetry to reduce noise and cost. We are a 100% remote company with 1,400+ team members across 40+ countries, and we’re backed by leading investors including Lightspeed Venture Partners, Sequoia Capital, GIC, Coatue, J.P. Morgan, CapitalG, and Lead Edge Capital.