
Senior Site Reliability Engineer
Honeycomb.io17 days ago
Remote, United KingdomSenior
Responsibilities
- Scale backend systems to support Honeycomb’s highest-volume customers.
- Lead or contribute to cross-team projects that improve reliability, scalability, infrastructure efficiency, and developer experience.
- Collaborate with backend teams to improve the use and performance of infrastructure.
- Become and serve as an Incident Commander, and train others in the role.
- Participate in the EU side of a new follow-the-sun on-call rotation.
- Help the organization navigate tradeoffs between reliability and other goals.
- Support a healthy cross-Atlantic engineering culture through transparent communication and feedback.
- Optionally represent Honeycomb through blog posts, conference talks, and presentations.
Requirements
- Strong experience with AWS and Kubernetes.
- Experience performing cost analysis and cost reduction.
- Solid experience with Helm, Terraform, and CI/CD.
- Project management skills.
- Software engineering experience; Golang and performance engineering experience are advantages.
- Experience with Kafka or another high-volume distributed system.
- Excellent written and spoken communication skills, including adapting communication to the audience and providing direct feedback.
- Familiarity with observability concepts such as SLOs and instrumentation, and with data-driven decision-making.
- Comfort operating in ambiguity with a bias toward action and experimentation.
- Interest in both the technical and human aspects of reliability engineering.
- Experience working in geographically distributed teams.
Benefits
- Fully distributed, remote-first work arrangement.
- Unlimited paid time off.
- Home office, co-working, and internet stipend.
- Full benefits coverage for employees, with additional dependent coverage available.
- Up to 16 weeks of paid parental leave regardless of path to parenthood.
- Annual development allowance.
- Generous equity with an employee-friendly stock program.
- Transparent pay levels based on experience.
Tech Stack
Categories
Site Reliability