14 days ago
Base Salary
$250k - $320k/yr
Responsibilities
- Write and optimize compute kernels for a custom AI accelerator, including tensor operations, data movement, and memory-hierarchy exploitation.
- Develop and maintain profiling infrastructure to measure kernel performance against architectural targets.
- Define and document shuffle patterns for ML kernel primitives across control, tensor-core, and CUTLASS-style operations.
- Drive kernel DSL decisions involving thread spawning, register passing, and memory management.
- Enable end-to-end kernel execution on the architectural simulator.
- Collaborate with compiler engineers on the MLIR dialect and use kernels as its primary validation target.
- Create onboarding documentation and kernel-writing guides.
Requirements
- Production-grade C/C++ systems programming for performance-critical kernels.
- Deep CUDA or equivalent accelerator-programming experience, including GPU kernels, warp or wavefront execution, memory coalescing, and shared-memory optimization.
- Strong computer architecture knowledge, including pipelines, memory hierarchies, data movement costs, and software-to-hardware mapping.
- Experience with performance profiling and optimization, including bottleneck identification, throughput measurement, and iterative tuning.
- Practical understanding of GEMM, convolution, attention, reduction, and scatter/gather tensor operations as mapped to hardware.
- Python experience for scripting, DSL integration, and profiling automation.
- Optional experience with RISC-V, x86, ARM64, MLIR, LLVM, HPC, scientific computing, FPGA, Verilog/SystemVerilog, CUTLASS, Triton, or similar kernel libraries.
Benefits
- Equity grant according to company guidelines.
- Medical, dental, and vision benefits.
- 401(k) plan.
- Standard paid time off.
- California pay range of $250,000–$320,000 USD.
- DensityAI provides immigration support for qualified candidates and requires work authorization at the start of employment.
