Base Salary
$175k - $210k/yr
Responsibilities
- Define technical strategy, architectural vision, and roadmap priorities for Adaptive Telemetry.
- Lead planning, design, execution, rollout, and long-term operation of large cross-functional initiatives.
- Own architecture, reliability, performance, scalability, availability, latency, maintainability, and cost for critical systems.
- Define SLOs and SLIs, lead high-severity incident response, run blameless postmortems, and drive systemic fixes and automation.
- Improve observability, alerting, runbooks, capacity planning, and operational readiness to reduce toil and MTTR.
- Coordinate with Product, Design, and other teams to align priorities, negotiate tradeoffs, and remove blockers.
- Mentor senior and mid-level engineers, lead design reviews, and raise engineering standards.
- Communicate technical strategy to non-engineering stakeholders and represent engineering in cross-team planning.
Requirements
- Proven experience delivering and operating complex distributed systems spanning multiple teams.
- Deep systems-design understanding across latency, consistency, availability, scalability, and cost tradeoffs.
- Hands-on experience with cloud-native architectures, containers, Kubernetes, infrastructure as code, and platform operations.
- Experience defining SLOs and SLIs, capacity planning, performance tuning, and end-to-end reliability work.
- Strong coding and technical design skills; experience with Go is used, while Python, C, C++, Rust, or similar languages are transferable.
- Curiosity and comfort with AI-assisted development tools, ideally including practical experience integrating them into team workflows.
- Familiarity with streaming or messaging systems such as Kafka and observability tooling such as Prometheus and Grafana.
- Ability to influence without authority, align cross-functional stakeholders, set priorities, and drive outcomes remotely.
- Clear written and verbal communication across technical and non-technical audiences.
- Self-directed, customer-focused approach with a bias toward action.
Benefits
- 100% remote work for candidates in USA time zones.
- In-person onboarding is provided.
- 30 days of annual leave per year, including 3 Grafana Shutdown Days, subject to local legislation.
- Equity and bonus benefits may be available.
- Career growth pathways and a transparent, high-trust company culture.
Tech Stack
About Grafana
Grafana Labs, the company behind the open observability cloud, is founded on the principles of open source, open standards, open ecosystems, and open culture. Grafana Cloud, our fully managed observability platform, is flexible and built for scale. With Grafana Cloud's actually useful AI, organizations can see, understand, and act on all their disparate data to move at the speed of their ambitions, while getting the visibility they need to run AI systems reliably and at scale. Today, more than 35 million users and 7,000+ customers – including Anthropic, Bloomberg, NVIDIA, Microsoft, and Salesforce – trust Grafana Labs to ensure reliability of their applications and systems, resolve incidents quickly, and optimize their telemetry to reduce noise and cost. We are a 100% remote company with 1,400+ team members across 40+ countries, and we’re backed by leading investors including Lightspeed Venture Partners, Sequoia Capital, GIC, Coatue, J.P. Morgan, CapitalG, and Lead Edge Capital.