2 months ago
Bengaluru, IndiaStaff+
Responsibilities
- Design and optimize neural-network operators for performance-critical workloads.
- Develop new neural-network operators using CUDA and custom runtime APIs.
- Drive runtime-level optimizations across compute, memory, and scheduling.
- Own runtime-to-neural-network interfaces and the execution model.
- Implement and optimize operator fusion, including examples such as matrix multiplication, bias, and LayerNorm.
- Identify and resolve performance bottlenecks across the software stack.
- Collaborate with compiler, PyTorch framework, and low-level software teams.
- Drive performance, scalability, and hardware utilization for AI workloads on the accelerator platform.
Requirements
- At least 8 years of experience in systems software, runtime, or performance engineering.
- Strong experience with PyTorch, TensorFlow, JAX, or similar frameworks.
- Experience developing and optimizing neural-network operators or kernels, operator fusion, and graph-level optimizations.
- Hands-on expertise in C/C++ and CUDA or similar low-level programming.
- Experience with runtime systems and execution engines.
- Strong understanding of memory hierarchy, data movement, and parallel execution.
- Preferred experience with GPUs, NPUs, or custom AI accelerators.
- Preferred familiarity with XLA, MLIR, and compiler-runtime interaction.
- Preferred experience optimizing large language model or large-scale deep-learning workloads.
Benefits
- Equal-opportunity and inclusive employment environment with accommodation support for applicants with disabilities.
- The anticipated application deadline is May 6, 2026, subject to earlier closure if the position is filled.