1 day ago
San Jose, CA, USASenior
Responsibilities
- Design and develop distributed communication and execution capabilities in the Triton AMDGPU backend for scalable multi-GPU AI workloads.
- Implement compiler and runtime mechanisms for GPU-initiated communication, collective operations, remote memory access, synchronization, and distributed execution.
- Optimize compute, communication, inter-GPU data movement, communication/computation overlap, memory hierarchy utilization, and GPU-driven scheduling.
- Develop and optimize distributed Triton kernels and execution models for high performance and scalability.
- Profile, debug, and resolve cross-stack issues spanning the Triton compiler, runtime, ROCm, and GPU hardware.
- Collaborate with GPU architecture, compiler, runtime, and performance teams on distributed GPU programming capabilities.
- Contribute to the open-source Triton and ROCm distributed ecosystems.
Requirements
- Deep experience in compiler development, GPU software, distributed systems, or performance engineering.
- Familiarity or hands-on experience with the Triton compiler and runtime.
- Deep understanding of GPU execution models, memory hierarchy, scheduling, occupancy, and performance characteristics.
- Understanding of GPU runtime systems, communication stacks, and multi-GPU interconnects including XGMI, NVLink, PCIe, or InfiniBand.
- Familiarity with distributed GPU communication libraries including RCCL, NCCL, NVSHMEM, rocSHMEM, or MPI.
- Experience developing, optimizing, and scaling workloads across multiple GPUs, including inter-GPU communication and synchronization.
- Strong GPU programming experience with Triton, HIP, CUDA, or similar parallel programming environments.
- Strong knowledge of MLIR and/or LLVM internals.
- Experience profiling, debugging, and optimizing across compiler, runtime, and hardware layers.
- Familiarity with ROCm, HIP, CUDA, and GPU performance profiling and optimization tools.
- Experience optimizing large-scale AI, machine learning, or HPC workloads across multi-GPU systems.
- Experience contributing to open-source projects and working in collaborative engineering environments.
- Bachelor’s or master’s degree in Computer Engineering, Computer Science, or Electrical Engineering, or equivalent practical experience.
- Strong problem-solving, communication, and technical leadership skills.
Benefits
- Hybrid role based in San Jose, California.
- AMD benefits are offered.
- This role is not eligible for visa sponsorship.
About AMD
We care deeply about transforming lives with AMD technology to enrich our industry, our communities, and the world. Our mission is to build great products that accelerate next-generation computing experiences – the building blocks for the data center, artificial intelligence, PCs, gaming and embedded. Underpinning our mission is the AMD culture. We push the limits of innovation to solve the world’s most important challenges. We strive for execution excellence while being direct, humble, collaborative, and inclusive of diverse perspectives. AMD together we advance_
