16 days ago
Remote, Czechia or Prague, CzechiaSenior
Responsibilities
- Design, build, and operate metrics, logging, distributed tracing, and alerting platform components for large-scale infrastructure.
- Build high-volume telemetry pipelines while managing cardinality, retention, cost, and query-performance tradeoffs.
- Define SLO/SLI frameworks and low-noise alerting strategies for on-call engineers.
- Partner with service delivery and operations teams to build incident-focused observability capabilities.
- Integrate observability tooling with incident-management workflows, root-cause analysis, and post-incident reviews.
- Improve detection and resolution performance, including MTTD and MTTR.
- Contribute to AI-assisted operations tooling such as automated triage, anomaly detection, and engineer-assist tools.
- Own the reliability, scalability, and security of the observability stack.
- Document architecture, runbooks, and operational practices.
Requirements
- Proven experience designing and building observability platforms for large-scale production infrastructure rather than only consuming existing tooling.
- Strong hands-on experience with metrics, logging, and distributed tracing tools such as Prometheus, Grafana, OpenTelemetry, Loki, Thanos, Cortex, Mimir, Elasticsearch, OpenSearch, Jaeger, or Tempo.
- Experience with high-volume telemetry pipelines and tradeoffs involving cardinality, retention, cost, and query latency.
- Strong software engineering skills in at least one relevant language, such as Go, Python, or Rust.
- Experience with Kubernetes and cloud-native infrastructure.
- Strong understanding of SLO/SLI and error-budget practices and alerting design.
- Ability to work directly with operations and service delivery teams in a fast-moving infrastructure environment.
- Preferred: observability experience for GPU/HPC or other high-performance compute environments.
- Preferred: experience with eBPF-based observability tooling.
- Preferred: familiarity with AIOps, ML-based anomaly detection, or automated triage systems.
- Preferred: managed-services or MSP experience tied to customer-facing SLAs.
- Preferred: contributions to open-source observability projects.
Benefits
- Professional development and training.
- Conference and working-group attendance.
- Company outings, happy hours, hackathons, and tech talks.
- Competitive compensation package with a strong benefits plan.
- Work with an established cloud-infrastructure company, passionate colleagues, Fortune 500 and Global 2000 customers, and open-source technologies.
- The posting does not state a work arrangement or contract duration.
