Software Engineer – GPU Kernel
FriendliAI5 months ago
Seoul, Korea, SouthMid Level
Responsibilities
- Design, implement, and optimize high-performance GPU kernels for AI inference, including GEMM, attention, and routing.
- Develop and maintain GPU code in CUDA and C++, including low-level assembly when needed.
- Implement reduced-precision and quantized FP8 and FP4 kernels for low-latency and high-throughput inference.
- Benchmark NVIDIA and AMD hardware and ensure cross-vendor performance parity.
- Contribute to internal GPU libraries and tune performance-critical components.
- Accelerate multimodal model pipelines.
- Investigate and integrate next-generation GPU features.
- Collaborate with the platform team to deploy optimized inference work into production.
Requirements
- 3+ years of experience in GPU programming, HPC, or performance-critical systems.
- Bachelor’s or Master’s degree in Computer Science, Computer Engineering, Electrical Engineering, or a related field.
- Strong proficiency in CUDA for NVIDIA GPUs or ROCm/HIP for AMD GPUs.
- Deep understanding of GPU architecture, including warps, threads, memory hierarchy, synchronization, and latency-throughput trade-offs.
- Proficiency in C++.
- Experience with GPU profiling and performance tuning.
- Strong numerical background, including precision trade-offs and quantization techniques.
- Preferred experience optimizing transformer, multimodal, or Mixture-of-Experts architectures at the kernel level.
- Familiarity with GPU libraries and frameworks such as CUTLASS and Triton.
- Experience with inter-GPU communication programming.
- Open-source contributions related to GPU performance or machine-learning acceleration.
- Research or conference presentations on GPU optimization, HPC, or numerical computing.
Benefits
- Flexible working hours.
- Daily lunch and dinner, plus unlimited snacks and beverages.
- Supportive and highly collaborative work environment.
- Health check-up support and top-tier equipment and hardware support.
- Competitive compensation, startup equity, and health insurance.
- The posting describes a small, fast-moving team working on generative AI infrastructure.