KRAFTON

[AI Research Div.] ML Infrastructure Engineer (5년 이상)

KRAFTON
Apply
8 days ago
Seoul, Korea, SouthSenior

Responsibilities

  • Operate, stabilize, improve, and automate the B300 125-node GPU infrastructure.
  • Design, build, and operate Kubernetes-based ML/GPU platform capabilities including scheduling, multi-tenancy, workload isolation, quotas, observability, and incident response.
  • Develop operational strategies to improve GPU utilization, workload wait times, throughput, and cost efficiency.
  • Continuously improve the ML platform and reproducible operating processes for research and development teams.
  • Coordinate requirements across research and development organizations and make platform-level technical decisions.
  • Analyze system-wide incidents and performance issues, identify root causes, and implement structural improvements.

Requirements

  • Hands-on experience improving ML/GPU platform scheduling, workload isolation, observability, and incident response systems.
  • Experience analyzing failures and performance issues from a system-wide perspective and implementing root-cause and structural improvements.
  • Experience defining common ML/GPU infrastructure requirements with research and development teams and proposing technical alternatives and priorities.
  • Experience using generative AI, LLM-based tools, or code assistants to improve operational efficiency, problem solving, documentation, or automation productivity.
  • No restriction preventing overseas business travel.
  • Preferred experience operating and optimizing next-generation GPU architectures such as B200, B300, H100, H200, GB200, or GB300.
  • Preferred experience optimizing GPU communication performance in NCCL, RDMA, RoCE, or InfiniBand environments.
  • Preferred experience operating and optimizing distributed storage such as Ceph or MinIO for AI workloads.
  • Preferred experience applying GPU resource allocation, scheduling, priority, and cost-performance optimization strategies in production.
  • Preferred experience with NVIDIA GPU Operator, DCGM, MIG, MPS, Run:ai, Slurm, Kueue, or Volcano.
  • Preferred experience improving distributed training throughput and GPU utilization by addressing data loading, communication bottlenecks, or training scheduling.

Benefits

  • Permanent employment with no employment-type or salary adjustment during the probationary period.
  • Five-month probationary period, which may end early or result in non-continuation based on evaluation.
  • Work location: Yeoksam Centerfield West Tower.
  • Employment preference is provided to eligible individuals with disabilities and distinguished service recipients according to applicable law.

Tech Stack

KRAFTON

About KRAFTON

1,001-5,000 employees
Contact me