Responsibilities
- Own production reliability for high-SLA and complex customer environments supporting Grafana Cloud databases based on Mimir, Loki, Tempo, and Pyroscope.
- Define and evolve per-tenant SLOs and reliability models while reducing SLO budget burn and preventing repeat incidents.
- Design and implement automation, self-healing, auto-scaling, monitoring improvements, and other solutions that reduce operational toil.
- Improve customer observability, reliability, scalability, fault tolerance, and production operability across cloud environments.
- Serve as a primary escalation point and participate in on-call and customer-impacting incident response through resolution, post-incident review, and customer communication.
- Lead or contribute to incident response, bridge calls, PIRs, design documents, code reviews, and technical reviews.
- Influence feature design, product strategy, roadmaps, and technical designs to support production scalability and reliability.
- Work as an embedded SRE partner within the Mimir, Loki, and Tempo engineering squads.
- Teach Site Reliability Engineering practices and communicate reliability best practices to engineering teams.
- Provide technical leadership, mentorship, and regular collaboration with managers, colleagues, engineering leaders, and product engineers.
Requirements
- At least 8 years of engineering experience, including 4 or more years in SRE, customer reliability engineering, or production engineering.
- Strong Kubernetes experience in AWS, GCP, or Azure and familiarity with Helm, Terraform, Jsonnet, or similar infrastructure-as-code tooling.
- Experience operating multi-tenant production systems and designing and implementing SLOs.
- Experience with one or more programming languages such as Go, Python, or Java.
- Experience with Linux operating system internals and knowledge of networking, cloud storage, and scaling.
- Demonstrated technical leadership, project leadership, mentoring, and ability to act as a force multiplier.
- Experience participating in incident response, producing post-incident reviews, and following through on corrective actions.
- Ability to reason about performance, scaling, failure modes, troubleshooting, and reliability tradeoffs.
- Ability to partner deeply with product engineering teams and work autonomously in a transparent, collaborative environment.
- Formal customer reliability engineering experience is strongly preferred.
Benefits
- 100% remote role for candidates in the UK, Sweden, Spain, Germany, or Ireland.
- In-person onboarding is provided.
- 30 days of annual leave per year, including 3 Grafana Shutdown Days, subject to local legislation.
- Equity, bonus where applicable, and other company benefits are included.
- Career growth pathways, transparent communication, autonomous teams, and a global collaborative culture.
Categories
About Grafana
Grafana Labs, the company behind the open observability cloud, is founded on the principles of open source, open standards, open ecosystems, and open culture. Grafana Cloud, our fully managed observability platform, is flexible and built for scale. With Grafana Cloud's actually useful AI, organizations can see, understand, and act on all their disparate data to move at the speed of their ambitions, while getting the visibility they need to run AI systems reliably and at scale. Today, more than 35 million users and 7,000+ customers – including Anthropic, Bloomberg, NVIDIA, Microsoft, and Salesforce – trust Grafana Labs to ensure reliability of their applications and systems, resolve incidents quickly, and optimize their telemetry to reduce noise and cost. We are a 100% remote company with 1,400+ team members across 40+ countries, and we’re backed by leading investors including Lightspeed Venture Partners, Sequoia Capital, GIC, Coatue, J.P. Morgan, CapitalG, and Lead Edge Capital.