3 months ago
Base Salary
$175k - $210k/yr
Responsibilities
- Define technical strategy, architectural vision, and roadmap priorities for Adaptive Telemetry.
- Lead planning, design, execution, rollout, and long-term operation of large cross-functional initiatives.
- Own architecture, reliability, performance, scalability, availability, latency, maintainability, and cost for critical systems.
- Define SLOs and SLIs, lead high-severity incident response, run blameless postmortems, and drive systemic fixes and automation.
- Improve observability, alerting, runbooks, capacity planning, and operational readiness to reduce toil and MTTR.
- Coordinate with Product, Design, and other teams to align priorities, negotiate tradeoffs, and remove blockers.
- Mentor senior and mid-level engineers, lead design reviews, and raise engineering standards.
- Communicate technical strategy to non-engineering stakeholders and represent engineering in cross-team planning.
Requirements
- Proven experience delivering and operating complex distributed systems spanning multiple teams.
- Deep systems-design understanding across latency, consistency, availability, scalability, and cost tradeoffs.
- Hands-on experience with cloud-native architectures, containers, Kubernetes, infrastructure as code, and platform operations.
- Experience defining SLOs and SLIs, capacity planning, performance tuning, and end-to-end reliability work.
- Strong coding and technical design skills; experience with Go is used, while Python, C, C++, Rust, or similar languages are transferable.
- Curiosity and comfort with AI-assisted development tools, ideally including practical experience integrating them into team workflows.
- Familiarity with streaming or messaging systems such as Kafka and observability tooling such as Prometheus and Grafana.
- Ability to influence without authority, align cross-functional stakeholders, set priorities, and drive outcomes remotely.
- Clear written and verbal communication across technical and non-technical audiences.
- Self-directed, customer-focused approach with a bias toward action.
Benefits
- 100% remote work for candidates in USA time zones.
- In-person onboarding is provided.
- 30 days of annual leave per year, including 3 Grafana Shutdown Days, subject to local legislation.
- Equity and bonus benefits may be available.
- Career growth pathways and a transparent, high-trust company culture.
Tech Stack
About Grafana
Grafana Labs builds open-source observability tools and a managed SaaS, Grafana Cloud, used by engineering and SRE teams to monitor, visualize, and analyze telemetry. Its products include Grafana dashboards plus Loki (logs), Tempo (traces), and Mimir (metrics), offered as cloud subscriptions and enterprise software. Privately held and headquartered in New York, it serves customers such as Anthropic, Bloomberg, NVIDIA, Microsoft, and Salesforce.
