8 hours ago
Remote, Spain +8 moreMid Level
Responsibilities
- Review and evaluate NKI kernel correctness, idiomatic Trainium development patterns, and implementation quality.
- Assess CUDA-to-NKI kernel migrations and cross-platform numerical correctness between CUDA, Triton, and NKI.
- Optimize and benchmark Trainium workloads, including tile-based computation, DMA scheduling, memory hierarchy usage, and hardware-specific bottlenecks.
- Analyze SBUF, PSUM, HBM, partition-dimension constraints, and other Trainium kernel optimization concerns.
- Provide clear written technical feedback and quality assessments of kernel implementations.
Requirements
- At least 2 years of hands-on experience developing or optimizing kernels with the Neuron Kernel Interface.
- Experience working with AWS Trainium and/or Inferentia2 hardware.
- Strong understanding of tile-based computation, SBUF/PSUM/HBM memory hierarchy, partition-dimension constraints, DMA orchestration, and Trainium-specific optimization.
- Ability to evaluate CUDA-to-NKI migrations and experience profiling and optimizing Trainium workloads.
- Understanding of numerical differences across GPU and Trainium backends.
- Strong ability to analyze complex technical implementations and provide clear written feedback.
- Preferred qualifications include experience with the AWS Neuron SDK or Neuron Compiler, CUDA or Triton kernel development, NeuronCore-v2, FP32/BF16/FP8/INT8 workloads, Trn1 or Trn2 benchmarking, nki.language, @nki.jit, XLA custom calls, technical evaluation, AI/ML data projects, RLHF, or rubric-based assessment.
Benefits
- Remote work arrangement.
- Part-time, project-based consulting engagement.