Software Engineer – GPU Kernel
FriendliAI6 months ago
San Francisco, CA, USAMid Level
Responsibilities
- Design, implement, and optimize high-performance GPU kernels for AI inference, including GEMM, attention, and routing.
- Develop and maintain GPU code in CUDA and C++, including low-level assembly when needed.
- Implement reduced-precision and quantized FP8/FP4 kernels for low-latency and high-throughput inference.
- Benchmark NVIDIA and AMD hardware and ensure cross-vendor performance parity.
- Contribute to internal GPU libraries and tune performance-critical components.
- Accelerate multimodal model pipelines and integrate next-generation GPU features.
- Collaborate with the platform team to deploy GPU kernel work into production.
Requirements
- At least 3 years of experience in GPU programming, HPC, or performance-critical systems.
- Bachelor’s or Master’s degree in Computer Science, Computer Engineering, Electrical Engineering, or a related field.
- Strong proficiency in CUDA for NVIDIA GPUs or ROCm/HIP for AMD GPUs.
- Deep understanding of GPU architecture, including warps, threads, memory hierarchy, synchronization, and latency-throughput trade-offs.
- Proficiency in C++.
- Experience with GPU profiling and performance tuning.
- Strong numerical background and understanding of precision trade-offs and quantization techniques.
- Preferred: experience optimizing transformer, multimodal, or Mixture-of-Experts architectures at the kernel level.
- Preferred: familiarity with CUTLASS and Triton.
- Preferred: inter-GPU communication programming experience, open-source GPU performance or ML acceleration contributions, and research or conference presentations on GPU optimization, HPC, or numerical computing.
Benefits
- Flexible working hours.
- Daily lunch and dinner, plus unlimited snacks and beverages.
- Supportive and highly collaborative work environment.
- Health check-up support and top-tier equipment and hardware support.
- Competitive compensation, startup equity, health insurance, and other benefits.