Mirantis

Observability Platform Engineer — Neocloud

Mirantis
Apply
6 days ago
Remote, EMEASenior

Responsibilities

  • Design, build, and operate metrics, logging, distributed tracing, and alerting platform components for large-scale infrastructure.
  • Build high-volume, high-cardinality telemetry pipelines while managing retention, cost, and query-performance tradeoffs.
  • Define and implement SLO/SLI frameworks, error-budget practices, and low-noise alerting strategies.
  • Partner with service delivery and operations teams to build incident-focused observability capabilities.
  • Integrate observability tooling with incident-management workflows, root-cause analysis, and post-incident reviews.
  • Improve incident detection and resolution speed, including MTTD and MTTR.
  • Contribute to AI-assisted operations capabilities such as automated triage, anomaly detection, and engineer-assist tooling.
  • Own the reliability, scalability, and security of the observability stack.
  • Document architecture, runbooks, and operational practices.

Requirements

  • Proven experience designing and building observability platforms for large-scale production infrastructure rather than only consuming existing systems.
  • Strong hands-on experience with metrics, logging, and distributed tracing tools such as Prometheus, Grafana, OpenTelemetry, Loki, Thanos, Cortex, Mimir, Elasticsearch, OpenSearch, Jaeger, or Tempo.
  • Experience building high-volume telemetry pipelines and managing cardinality, retention, cost, and query-latency tradeoffs.
  • Strong software engineering skills in at least one relevant language, such as Go, Python, or Rust.
  • Experience with Kubernetes and cloud-native infrastructure.
  • Strong understanding of SLO, SLI, error-budget, and alerting practices.
  • Ability to work directly with operations and service delivery teams in a fast-moving infrastructure environment.
  • Preferred experience with GPU/HPC or other high-performance compute observability.
  • Preferred experience with eBPF-based observability tooling.
  • Preferred familiarity with AIOps, ML-based anomaly detection, or automated triage systems.
  • Preferred experience in managed services or MSP environments with customer-facing SLAs.
  • Open-source observability contributions are preferred.

Benefits

  • Professional development and training.
  • Conference and working-group attendance.
  • Company outings, happy hours, hackathons, and tech talks.
  • Competitive compensation package with a strong benefits plan.
  • Opportunity to work with open-source cloud-infrastructure technologies and colleagues serving Fortune 500 and Global 2000 customers.

Tech Stack

ElasticsearchGoGrafanaKubernetesPrometheusPythonRust

Categories

Mirantis

About Mirantis

501-1,000 employees
Contact me