Anyone AI

GPU Kernel Engineer – CUDA, Triton & Accelerator Performance

Anyone AI
Apply
8 hours ago
Remote, Spain +8 moreMid Level

Responsibilities

  • Review GPU and accelerator kernel implementations for correctness and optimization.
  • Implement, debug, translate, migrate, fuse, and optimize kernels across frameworks and hardware platforms.
  • Compare outputs with reference implementations and evaluate numerical tolerance thresholds.
  • Profile kernels, review benchmarks, identify bottlenecks, and determine whether performance comparisons and targets are realistic.
  • Investigate compilation, driver, memory, shape, and runtime issues and distinguish defects from environment or optimization challenges.
  • Provide clear, actionable technical feedback on kernel implementations and technical tasks.

Requirements

  • At least 3 years of hands-on experience developing, optimizing, or debugging GPU or accelerator kernels.
  • Strong hands-on experience with at least two of CUDA, Triton, NKI/AWS Neuron, and Pallas/JAX.
  • Strong understanding of GPU performance optimization, including memory bandwidth, compute throughput, GPU occupancy, shared memory, register pressure, memory coalescing, and bank conflicts.
  • Experience with profiling tools such as Nsight, NCU, roofline analysis, or framework-native profilers.
  • Strong understanding of floating-point numerical correctness and tolerance thresholds.
  • Experience debugging kernel compilation and runtime issues.
  • Preferred qualifications include experience across NVIDIA GPU and custom accelerator ecosystems, AWS Trainium, TPU, JAX, compiler engineering, MLIR, XLA, GPU or ML kernel libraries, cuBLAS, cuDNN, Triton community kernels, JAX/XLA custom calls, AI model evaluation, RLHF, or technical benchmark development.

Benefits

  • Remote work arrangement.
  • Part-time, project-based consulting engagement.

Categories

Contact me