Responsibilities
- Design, build, and operate reconciliation systems, including the Stack State Service backend, to detect and repair configuration drift.
- Develop reliable lifecycle workflows across Stack State Service, grafana.com, Hosted Grafana, deployment configurations, cloud regions, and clusters.
- Manage rollouts for plugins, dashboards, data sources, Grafana versions, release channels, and stack-level configuration.
- Support new region and cluster rollouts and improve incident response, recovery paths, runbooks, dashboards, alerts, and rollout controls.
- Implement backend services and systems, write maintainable and tested code, and debug across service boundaries.
- Collaborate with Product, UX, Hosted Grafana, Infrastructure, Support, and adjacent AppCore squads.
- Contribute to roadmap planning, technical design, prioritization, on-call improvements, and long-term simplification of stack operations.
- Participate in customer feedback response and a follow-the-sun on-call rotation when ready.
Requirements
- At least 1 year of fully remote work experience.
- Experience working on a SaaS platform and familiarity with distributed-systems concepts such as scalability, multi-tenancy, and high availability.
- Professional experience with Golang and willingness to work across backend service and application code.
- Experience contributing to projects from initial brainstorming through customer delivery.
- Ability to write clean, well-tested, maintainable software and execute well-defined work iteratively.
- Willingness to collaborate across teams and align work with internal and external stakeholders.
- Familiarity with Kubernetes in AWS, GCP, or Azure and exposure to infrastructure-as-code tooling such as Helm, Terraform, or Jsonnet.
- Experience participating in blameless incident response and post-incident reviews.
- Preferred experience with TypeScript and Node.js.
- Preferred experience with Kubernetes control-plane patterns, operators, reconcilers, or desired-state systems.
- Preferred experience with Jsonnet, Tanka, Terraform, Flux, Argo, or similar deployment and configuration tooling.
- Preferred experience with SaaS provisioning, tenancy, regional expansion, plugin rollout, or customer lifecycle systems.
- Preferred experience handling configuration drift, partial failure, or cross-service state mismatch during incident response.
Benefits
- 100% remote work with a global, asynchronous culture.
- In-person onboarding.
- 30 days of annual leave per year, including 3 Grafana Shutdown Days, subject to local legislation.
- Restricted Stock Units are included with the role.
- Career growth pathways, transparent communication, empowered teams, and open-source collaboration.
- The role is available to candidates located in the UK, Germany, Spain, Ireland, and Sweden.
Tech Stack
Categories
About Grafana
Grafana Labs, the company behind the open observability cloud, is founded on the principles of open source, open standards, open ecosystems, and open culture. Grafana Cloud, our fully managed observability platform, is flexible and built for scale. With Grafana Cloud's actually useful AI, organizations can see, understand, and act on all their disparate data to move at the speed of their ambitions, while getting the visibility they need to run AI systems reliably and at scale. Today, more than 35 million users and 7,000+ customers – including Anthropic, Bloomberg, NVIDIA, Microsoft, and Salesforce – trust Grafana Labs to ensure reliability of their applications and systems, resolve incidents quickly, and optimize their telemetry to reduce noise and cost. We are a 100% remote company with 1,400+ team members across 40+ countries, and we’re backed by leading investors including Lightspeed Venture Partners, Sequoia Capital, GIC, Coatue, J.P. Morgan, CapitalG, and Lead Edge Capital.
