3 months ago
Remote, CanadaStaff+
Responsibilities
- Define technical strategy, architectural vision, and roadmap for adaptive telemetry systems.
- Lead planning, design, execution, rollout, and long-term operation of large cross-functional initiatives.
- Own architecture, reliability, performance, scalability, availability, latency, maintainability, and cost for critical systems.
- Define SLOs and SLIs, lead high-severity incident response, run blameless postmortems, and drive systemic fixes.
- Improve observability, alerting, runbooks, capacity planning, automation, and operational readiness.
- Coordinate with Product, Design, and other teams to align priorities, negotiate tradeoffs, and remove blockers.
- Mentor senior and mid-level engineers, lead design reviews, and raise engineering standards.
- Represent engineering strategy in internal and external cross-team planning.
Requirements
- Proven experience delivering and operating large, complex distributed systems spanning multiple teams.
- Deep understanding of latency, consistency, availability, scalability, and cost tradeoffs.
- Hands-on experience with cloud-native architectures, containers or Kubernetes, infrastructure as code, and operational practices.
- Experience defining SLOs and SLIs, capacity planning, tuning performance, and owning reliability initiatives.
- Strong coding and technical design skills; Grafana Labs uses Go, while Python, C, C++, Rust, or similar experience can translate.
- Curiosity and comfort with AI-assisted or agentic development tools, ideally with practical experience integrating them into team workflows.
- Familiarity with streaming or messaging systems such as Kafka and observability tooling such as Prometheus and Grafana.
- Ability to influence without authority, align cross-functional stakeholders, set priorities, and drive outcomes in a remote-first environment.
- Strong written and verbal communication skills for technical and non-technical audiences.
Benefits
- 100% remote position for candidates in Canadian time zones.
- In-person onboarding is provided.
- 30 days of annual leave per year, including 3 Grafana Shutdown Days, subject to local legislation.
- Career growth pathways, transparent communication, empowered teams, and a global collaborative culture.
- Company-funded usage budget for AI coding assistants.
Tech Stack
About Grafana
Grafana Labs builds open-source observability tools and a managed SaaS, Grafana Cloud, used by engineering and SRE teams to monitor, visualize, and analyze telemetry. Its products include Grafana dashboards plus Loki (logs), Tempo (traces), and Mimir (metrics), offered as cloud subscriptions and enterprise software. Privately held and headquartered in New York, it serves customers such as Anthropic, Bloomberg, NVIDIA, Microsoft, and Salesforce.
