about 12 hours ago
Base Salary
$154k - $185k/yr
Responsibilities
- Partner closely with product engineering squads to ensure reliability.
- Own production reliability for high-SLA customer environments.
- Design and implement automation to scale reliability practices.
- Define and evolve per-tenant SLOs and reliability models.
- Lead customer-impacting incident response and post-incident reviews.
- Contribute to design docs and code reviews.
- Build automation to eliminate toil and improve alert quality.
Requirements
- 6+ years of engineering experience, with 3+ in SRE/CRE/production engineering.
- Strong Kubernetes experience in AWS, GCP, or Azure.
- Experience operating multi-tenant systems in production.
- Strong experience designing and implementing SLOs.
- Proficiency in one or more programming languages (e.g., Go, Python, Java).
- Excellent problem-solving and troubleshooting skills.
- Ability to work autonomously within an engineering team.
Benefits
- 100% remote work with a global culture.
- 30 days of annual leave, including Grafana Shutdown Days.
- Career growth pathways and opportunities for development.
- Access to modern AI coding assistants and tools.
- Transparent communication and approachable leadership.