8 hours ago
Remote, Spain +8 moreMid Level
Responsibilities
- Review GPU and accelerator kernel implementations for correctness and optimization.
- Implement, debug, translate, migrate, fuse, and optimize kernels across frameworks and hardware platforms.
- Compare outputs with reference implementations and evaluate numerical tolerance thresholds.
- Profile kernels, review benchmarks, identify bottlenecks, and determine whether performance comparisons and targets are realistic.
- Investigate compilation, driver, memory, shape, and runtime issues and distinguish defects from environment or optimization challenges.
- Provide clear, actionable technical feedback on kernel implementations and technical tasks.
Requirements
- At least 3 years of hands-on experience developing, optimizing, or debugging GPU or accelerator kernels.
- Strong hands-on experience with at least two of CUDA, Triton, NKI/AWS Neuron, and Pallas/JAX.
- Strong understanding of GPU performance optimization, including memory bandwidth, compute throughput, GPU occupancy, shared memory, register pressure, memory coalescing, and bank conflicts.
- Experience with profiling tools such as Nsight, NCU, roofline analysis, or framework-native profilers.
- Strong understanding of floating-point numerical correctness and tolerance thresholds.
- Experience debugging kernel compilation and runtime issues.
- Preferred qualifications include experience across NVIDIA GPU and custom accelerator ecosystems, AWS Trainium, TPU, JAX, compiler engineering, MLIR, XLA, GPU or ML kernel libraries, cuBLAS, cuDNN, Triton community kernels, JAX/XLA custom calls, AI model evaluation, RLHF, or technical benchmark development.
Benefits
- Remote work arrangement.
- Part-time, project-based consulting engagement.