2 months ago
Remote, United StatesStaff+
Responsibilities
- Execute core observability and telemetry architecture projects integrating proprietary Go infrastructure with open-source tools.
- Design and optimize OpenTelemetry, Prometheus, and VictoriaMetrics workflows for high-volume storage-cluster metrics.
- Write clean, performant Go code for the software-defined storage control plane and provide rigorous code reviews.
- Participate in design, implementation, automated testing, usability reviews, and release activities within the Scrum lifecycle.
- Develop documentation and standards that help engineering teams instrument their components correctly.
- Participate in a global support and on-call rotation for the distributed storage platform.
Requirements
- At least 8 years of backend development experience.
- Deep proficiency in Go for building high-performance, low-overhead system components.
- Hands-on experience implementing and operating telemetry pipelines for clustered, distributed, or cloud-native solutions.
- Strong knowledge of Prometheus operators, alerting rules, and scraping mechanics, plus the VictoriaMetrics stack.
- Practical experience with OpenTelemetry, including custom collector configurations, instrumentation SDKs, and data processing.
- Understanding of Linux networking, filesystems, and clustered storage applications under heavy I/O workloads.
- Experience developing or extending Go components that ingest, process, and forward large streams of metrics, logs, and traces.
- Experience optimizing the CPU and memory footprint of monitoring agents.
- Experience writing robust unit and integration tests for telemetry components during live cluster upgrades.
- Ability to independently own complex technical initiatives, collaborate across distributed teams, conduct code reviews, and produce clear documentation.
Benefits
- Participation in a global team on-call rotation.
- Work within a geographically distributed engineering team.