6 days ago
Remote, Spain or Barcelona, SpainSenior
Responsibilities
- Design, build, and operate metrics, logging, distributed tracing, and alerting platform components for large-scale infrastructure.
- Build high-volume telemetry pipelines and manage tradeoffs involving cardinality, retention, cost, and query performance.
- Define SLO/SLI frameworks, error-budget practices, and low-noise alerting strategies.
- Partner with service delivery and operations teams to build incident-focused observability capabilities.
- Integrate observability tooling with incident-management workflows, root-cause analysis, and post-incident review data.
- Improve detection speed and reduce mean time to detect and mean time to resolve across the platform.
- Contribute to AI-assisted operations capabilities such as automated triage, anomaly detection, and engineer-assist tooling.
- Own the reliability, scalability, and security of the observability stack.
- Document architecture, runbooks, and operational practices.
- Contribute to the observability roadmap and help ensure the platform scales with infrastructure growth.
Requirements
- Proven experience designing and building observability platforms for large-scale production infrastructure.
- Strong hands-on experience with metrics, logging, and distributed tracing tooling such as Prometheus, Grafana, OpenTelemetry, Loki, Thanos, Cortex, Mimir, Elasticsearch, OpenSearch, Jaeger, or Tempo.
- Experience with high-volume telemetry pipelines and related cardinality, retention, cost, and query-latency tradeoffs.
- Strong software engineering skills in at least one relevant language, such as Go, Python, or Rust.
- Experience with Kubernetes and cloud-native infrastructure.
- Strong understanding of SLO, SLI, error-budget, and alerting practices.
- Ability to work in a fast-moving environment while infrastructure and its observability platform are built together.
- Strong communication skills and ability to work with operations and service delivery teams.
- Preferred experience with GPU/HPC or other specialized high-performance compute infrastructure.
- Preferred experience with eBPF-based observability tooling.
- Familiarity with AIOps, ML-based anomaly detection, or automated triage systems is preferred.
- Experience in managed services or MSP environments is preferred.
- Contributions to open-source observability projects are preferred.
Benefits
- Professional development and training.
- Opportunities to attend conferences and working groups.
- Company outings, happy hours, hackathons, and tech talks.
- Competitive compensation package with a strong benefits plan.
- Work with an established cloud infrastructure company and passionate colleagues serving Fortune 500 and Global 2000 customers.
- Opportunities to contribute to open-source innovation.
- The posting does not specify a work arrangement or employment duration.
