8 days ago
Bengaluru, IndiaStaff+
Responsibilities
- Architect high-performance inference runtimes, kernel dispatchers, and memory planners for diffusion and transformer workloads.
- Investigate cross-GPU performance bottlenecks, communication overheads, and scheduling inefficiencies.
- Drive model, pipeline, and tensor parallelization strategies across multiple GPUs.
- Establish GPU optimization standards, tooling, and service-level indicators across the company.
- Collaborate with research teams on scalable implementations of novel architectures.
- Mentor engineers in profiling, performance tuning, and low-level optimization.
- Partner with hardware vendors and infrastructure teams to maximize cluster utilization.
Requirements
- At least 5 years of experience in high-performance computing, GPU runtime systems, or ML infrastructure.
- Proven expertise in CUDA, Triton, and C++, with deep knowledge of GPU scheduling, occupancy, register usage, and tensor cores.
- Experience building and maintaining distributed inference or training systems.
- Ability to design abstractions that balance flexibility and performance.
- Strong knowledge of NCCL, NVLink, PCIe, and interconnects.
- Familiarity with profiling automation and performance dashboards.
- Excellent technical leadership and mentoring capabilities.
- Preferred: experience with compiler-aided optimization using TVM, XLA, MLIR, or Triton.
- Preferred: experience tuning Stable Diffusion or transformer inference pipelines.
- Preferred: exposure to heterogeneous compute backends including AMD ROCm, TPU, or ASICs.
- Preferred: experience with hardware-software co-design initiatives and open-source or research contributions in GPU optimization.
