1 hour ago
Bengaluru, IndiaSenior
Responsibilities
- Design and build Kubernetes Custom Resource Definitions and operators for ML workloads, GPU node pools, and RayService resources.
- Architect and operate Kubernetes networking layers, including Gateway API, service load balancers, and API gateways for hybrid-cloud connectivity.
- Design and operate multi-NIC Kubernetes clusters with RDMA-enabled networking using GPUDirect RDMA, RoCE, and InfiniBand.
- Build and optimize an AI-aware GPU scheduler with topology-aware placement, preemption, and GPU defragmentation.
- Manage GPU pools across on-premise and cloud-burst environments, including spot or preemptible GPU scheduling.
- Deploy and operate KubeRay infrastructure for distributed Ray training and inference workloads.
- Automate GPU provisioning, node configuration, and cluster lifecycle management using infrastructure-as-code and GitOps.
- Implement networking policies, multi-tenant isolation, RBAC, and Kubernetes security controls.
- Build observability for GPU utilization, NCCL communication, scheduler decisions, and network throughput.
- Collaborate with ML Platform, AI Research, and Networking teams to optimize training and online-inference infrastructure.
Requirements
- Bachelor’s or master’s degree in Computer Science, Engineering, or a related field.
- 5+ years of experience building distributed systems or infrastructure platforms with deep Kubernetes expertise.
- Strong programming skills in Go and/or Python.
- Experience with Kubernetes controller development frameworks such as Kubebuilder, Operator SDK, or controller-runtime.
- Deep understanding of Kubernetes internals, including the API server, etcd, scheduler, controller manager, and kubelet.
- Hands-on experience with Kubernetes networking, Gateway API, CNI plugins, service load balancing, and hybrid-cloud architectures.
- Experience operating multi-NIC Kubernetes clusters using NVIDIA Network Operator, SR-IOV device plugins, or equivalent tools.
- Strong understanding of RDMA protocols and NCCL configuration for distributed GPU workloads.
- Experience with Kubernetes GPU scheduler frameworks, GPU pool management, MIG partitioning, and preemption policies.
- Hands-on experience deploying and operating KubeRay for distributed Ray workloads.
- Experience with GPU asset lifecycle management, bare-metal provisioning automation, and GitOps-based CD tooling such as Argo CD.
- Familiarity with NVIDIA CUDA, NVLink, NVSwitch, device plugins, and the GPU Operator ecosystem.
- Experience with Prometheus, Grafana, and OpenTelemetry.
- Strong debugging and performance-optimization skills across GPU driver stacks, RDMA networking, and distributed Kubernetes infrastructure.
