8 days ago
Seoul, Korea, SouthSenior
Responsibilities
- Design, build, and operate production-scale ML serving architectures for real-time in-game LLM services, coding agents, and diverse research models.
- Establish common serving systems for advanced deep-learning models across the AI Research division.
- Build integrated monitoring, logging, alerting, and observability systems for large-scale serving environments.
- Design CI/CD and GitOps systems supporting automated testing, zero-downtime model updates, and automatic rollback.
- Research and apply inference acceleration and serving optimization techniques to reduce latency and serving costs.
- Improve KV cache offloading, loading, and sharing across GPU, CPU, and disk, along with continuous batching and PagedAttention.
- Analyze and resolve failures and latency issues across inference engines, frameworks, and GPU hardware.
- Make architecture-level technical decisions and evaluate trade-offs involving cost, latency, and reliability.
Requirements
- At least 5 years of experience designing, building, and operating AI/LLM model-serving systems in production environments with large-scale traffic.
- Hands-on experience applying vLLM or TensorRT-LLM to optimize and stabilize high-performance inference systems.
- Experience implementing KV cache offloading/loading, PagedAttention, quantization, or other memory and compute efficiency techniques.
- Strong experience diagnosing and improving inference performance, failures, and latency across inference engines, frameworks, and GPU hardware.
- Experience directly building and improving monitoring, observability, deployment automation, and CI/CD environments for serving reliability.
- Ability to travel internationally without restrictions.
- Ability to lead architecture-level decisions and clearly understand solution trade-offs quantitatively and qualitatively is preferred.
- Contributions to vLLM, LMCache, or other open-source serving ecosystems are preferred.
- Deep understanding of vLLM, TensorRT-LLM, PyTorch Dynamo, FlashInfer, and NVIDIA Triton is preferred.
- Experience developing custom kernels and working with low-level CUDA acceleration is preferred.
- Required application materials include an application form, personal statement, career description, transcript, and portfolio.
Benefits
- Permanent employment, subject to possible adjustment based on the candidate’s negotiated terms.
- Work location is KRAFTON Yeoksam Centerfield West Tower.
- A five-month probationary period applies, with no change to employment type or salary during the period.
- The recruitment process may include an assignment, personality assessment, technical-fit interview, two to three culture-fit interviews, and additional testing if needed.
- Applicants eligible for employment protection under applicable laws, including people with disabilities and national merit recipients, receive preference.
