about 2 hours ago
Responsibilities
- Define and lead Logging, Metrics, Tracing & Alerting strategies.
- Solve scaling bottlenecks in telemetry data pipelines.
- Create tooling to accelerate root-cause analysis and reduce MTTR.
- Guide and mentor teams on monitoring best practices.
- Identify and tackle architectural bottlenecks to improve reliability.
- Drive features and products from ideation to implementation.
- Collaborate on planning and refinement to ensure engineering alignment.
- Respond to alerts and provide support for production issues.
Requirements
- 5+ years of experience with highly distributed systems.
- Expertise in designing APIs and data pipelines for real-time data ingestion.
- Experience improving software reliability across various categories.
- Proficiency in Go and/or Java for software development.
- Advanced understanding of system design and tradeoffs.
- Strong stakeholder management and technical communication skills.
- Motivated to learn and develop continuously.
- Experience with containerization and orchestration technologies.
- Hands-on experience with core telemetry data stores at scale.