8 days ago
Seoul, Korea, SouthSenior
Responsibilities
- Operate, stabilize, improve, and automate the B300 125-node GPU infrastructure.
- Design, build, and operate Kubernetes-based ML/GPU platform capabilities including scheduling, multi-tenancy, workload isolation, quotas, observability, and incident response.
- Develop operational strategies to improve GPU utilization, workload wait times, throughput, and cost efficiency.
- Continuously improve the ML platform and reproducible operating processes for research and development teams.
- Coordinate requirements across research and development organizations and make platform-level technical decisions.
- Analyze system-wide incidents and performance issues, identify root causes, and implement structural improvements.
Requirements
- Hands-on experience improving ML/GPU platform scheduling, workload isolation, observability, and incident response systems.
- Experience analyzing failures and performance issues from a system-wide perspective and implementing root-cause and structural improvements.
- Experience defining common ML/GPU infrastructure requirements with research and development teams and proposing technical alternatives and priorities.
- Experience using generative AI, LLM-based tools, or code assistants to improve operational efficiency, problem solving, documentation, or automation productivity.
- No restriction preventing overseas business travel.
- Preferred experience operating and optimizing next-generation GPU architectures such as B200, B300, H100, H200, GB200, or GB300.
- Preferred experience optimizing GPU communication performance in NCCL, RDMA, RoCE, or InfiniBand environments.
- Preferred experience operating and optimizing distributed storage such as Ceph or MinIO for AI workloads.
- Preferred experience applying GPU resource allocation, scheduling, priority, and cost-performance optimization strategies in production.
- Preferred experience with NVIDIA GPU Operator, DCGM, MIG, MPS, Run:ai, Slurm, Kueue, or Volcano.
- Preferred experience improving distributed training throughput and GPU utilization by addressing data loading, communication bottlenecks, or training scheduling.
Benefits
- Permanent employment with no employment-type or salary adjustment during the probationary period.
- Five-month probationary period, which may end early or result in non-continuation based on evaluation.
- Work location: Yeoksam Centerfield West Tower.
- Employment preference is provided to eligible individuals with disabilities and distinguished service recipients according to applicable law.
