Coupang

Sr. Staff Observability Engineer (GPU Cloud & Telemetry Platform)

Coupang
Apply
3 months ago
Seoul, Korea, SouthStaff+

Responsibilities

  • Own the end-to-end observability platform and telemetry strategy for GPU infrastructure.
  • Architect scalable, low-latency metric, log, and trace pipelines for GPU workloads, Kubernetes, containers, and distributed systems.
  • Build and optimize Alloy-to-Mimir metric pipelines and Vector-to-Loki log pipelines, including routing, retention, partitioning, sampling, and scaling.
  • Establish observability for GPU hardware, MIG partitions, Kubernetes GPU operators, scheduling, datacenter systems, and high-performance networks.
  • Define GPU-specific SLIs and SLOs covering utilization, scheduling latency, fragmentation, thermal conditions, and power anomalies.
  • Create Grafana dashboards for GPU fleet health, tenant usage, billing insights, capacity planning, and forecasting.
  • Integrate observability with CI/CD and Terraform/Kubernetes infrastructure-as-code pipelines to support canary analysis and automated rollbacks.
  • Develop Go/Python automation for telemetry pipeline health monitoring, dynamic routing, and workload scaling.
  • Lead cross-layer debugging and root-cause analysis for GPU contention, ML workload degradation, and telemetry pipeline failures.
  • Mentor engineers, lead design reviews, promote observability-by-design, and drive collaboration across SRE, infrastructure, and ML platform teams.
  • Guide Grafana and OpenTelemetry ecosystem adoption, interoperability, and build-versus-buy decisions.
  • Design secure telemetry pipelines with encryption, multi-tenant isolation, RBAC, data residency compliance, and Zero Trust patterns.

Requirements

  • BS/MS in Computer Science or equivalent practical experience.
  • Extensive experience in observability, SRE, or distributed infrastructure.
  • Proven experience building large-scale metrics and log telemetry pipelines.
  • Experience with Grafana Alloy or the Prometheus ecosystem, Grafana Mimir or Cortex/Thanos, Grafana Loki, and Datadog Vector or similar log pipelines.
  • Strong programming ability in Go or Python.
  • Experience with time-series databases and log storage at scale.
  • Experience with Kubernetes and Linux internals.
  • Experience with GPU systems, including NVIDIA DCGM and the CUDA ecosystem.
  • Experience with bare-metal GPU clusters and hybrid cloud environments.
  • High-performance networking experience, with RDMA and InfiniBand preferred.

Benefits

  • Full-time regular position with a 12-week probation period, which may be skipped, shortened, or extended based on business needs.
  • Equal-opportunity workplace with applicable employment-protection considerations.

Tech Stack

DatadogGoGrafanaKubernetesLinuxPrometheusPythonTerraform

Categories

Coupang

About Coupang

5,001-10,000 employees

Coupang builds and operates a South Korea–focused e-commerce marketplace with an end-to-end logistics network (Rocket Delivery), plus food delivery, video streaming, and fintech under brands such as Coupang, Eats, and Play. Revenue comes from first-party retail, third-party marketplace services, advertising, and memberships (Rocket WOW). Founded in 2010, the company is headquartered in Seattle and is publicly listed on the NYSE (CPNG).

Contact me