3 months ago
Seoul, Korea, SouthStaff+
Responsibilities
- Own the end-to-end observability platform and telemetry strategy for GPU infrastructure.
- Architect scalable, low-latency metric, log, and trace pipelines for GPU workloads, Kubernetes, containers, and distributed systems.
- Build and optimize Alloy-to-Mimir metric pipelines and Vector-to-Loki log pipelines, including routing, retention, partitioning, sampling, and scaling.
- Establish observability for GPU hardware, MIG partitions, Kubernetes GPU operators, scheduling, datacenter systems, and high-performance networks.
- Define GPU-specific SLIs and SLOs covering utilization, scheduling latency, fragmentation, thermal conditions, and power anomalies.
- Create Grafana dashboards for GPU fleet health, tenant usage, billing insights, capacity planning, and forecasting.
- Integrate observability with CI/CD and Terraform/Kubernetes infrastructure-as-code pipelines to support canary analysis and automated rollbacks.
- Develop Go/Python automation for telemetry pipeline health monitoring, dynamic routing, and workload scaling.
- Lead cross-layer debugging and root-cause analysis for GPU contention, ML workload degradation, and telemetry pipeline failures.
- Mentor engineers, lead design reviews, promote observability-by-design, and drive collaboration across SRE, infrastructure, and ML platform teams.
- Guide Grafana and OpenTelemetry ecosystem adoption, interoperability, and build-versus-buy decisions.
- Design secure telemetry pipelines with encryption, multi-tenant isolation, RBAC, data residency compliance, and Zero Trust patterns.
Requirements
- BS/MS in Computer Science or equivalent practical experience.
- Extensive experience in observability, SRE, or distributed infrastructure.
- Proven experience building large-scale metrics and log telemetry pipelines.
- Experience with Grafana Alloy or the Prometheus ecosystem, Grafana Mimir or Cortex/Thanos, Grafana Loki, and Datadog Vector or similar log pipelines.
- Strong programming ability in Go or Python.
- Experience with time-series databases and log storage at scale.
- Experience with Kubernetes and Linux internals.
- Experience with GPU systems, including NVIDIA DCGM and the CUDA ecosystem.
- Experience with bare-metal GPU clusters and hybrid cloud environments.
- High-performance networking experience, with RDMA and InfiniBand preferred.
Benefits
- Full-time regular position with a 12-week probation period, which may be skipped, shortened, or extended based on business needs.
- Equal-opportunity workplace with applicable employment-protection considerations.
About Coupang
Coupang builds and operates a South Korea–focused e-commerce marketplace with an end-to-end logistics network (Rocket Delivery), plus food delivery, video streaming, and fintech under brands such as Coupang, Eats, and Play. Revenue comes from first-party retail, third-party marketplace services, advertising, and memberships (Rocket WOW). Founded in 2010, the company is headquartered in Seattle and is publicly listed on the NYSE (CPNG).
