1 day ago
Jakarta, IndonesiaSenior
Responsibilities
- Lead root-cause analysis, automate recurring operations, develop L1 and L2 runbooks, mentor engineers, and serve as technical authority during P0/P1 incidents.
- Define platform compatibility, upgrade, canary, rollback, observability, and GPU remediation standards.
- Architect and optimize local or cloud-based AI training and inference platforms using PyTorch and NVIDIA technologies.
- Investigate NCCL and distributed multi-GPU or multi-node performance across compute, network/fabric, and storage.
- Diagnose issues across Linux, NVIDIA drivers, CUDA, container runtimes, Kubernetes or Red Hat OpenShift, NVIDIA GPU Operator, DCGM, XIDs, and MIG.
Requirements
- Strong Python and Bash automation, observability engineering, and deep platform troubleshooting experience.
- Strong NCCL, multi-GPU or multi-node performance troubleshooting, platform lifecycle upgrade, canary, and rollback experience.
- Strong PyTorch expertise plus experience with one or more of vLLM, llama.cpp, TensorRT, Triton, NIM, NGC, or NeMo.
- Expertise in Linux, NVIDIA drivers, CUDA, container runtimes, Kubernetes, NVIDIA GPU Operator, DCGM, and multi-GPU scheduling and health.
- At least 5 years of experience building and optimizing production AI training or inference platforms with substantial operations responsibility.
- Preferred experience with Red Hat OpenShift, production Kubernetes, distributed GPU networking and storage dependencies, NVIDIA DGX systems, enterprise multi-GPU platforms, mentoring, operational standards, and critical-incident leadership.
Benefits
- Hybrid-friendly work culture.
- Be Well programs supporting financial, mental, physical, and social health.
- Career-path tools, personalized development goals, continuous feedback, learning opportunities, certifications with Microsoft, Google, and Amazon, coaching, and hands-on experiences.
Tech Stack
Categories
About Kyndryl
Kyndryl is a public IT services company that designs, runs, and modernizes mission-critical infrastructure for large enterprises and governments. Spun off from IBM in 2021, it is headquartered in New York City and trades on the NYSE (ticker: KD). The company provides managed infrastructure, cloud migration and operations across AWS, Azure, and Google Cloud, plus mainframe, network, security, data, and digital workplace services.
