1 year ago
Responsibilities
- Write high-performance GPU kernels for novel model architectures.
- Integrate kernels into PyTorch pipelines through custom operations, extensions, dispatch, and benchmarking.
- Profile and optimize training and inference workflows to eliminate bottlenecks.
- Build correctness tests and numerical checks.
- Build and maintain performance benchmarks and guardrails to prevent regressions.
- Collaborate with researchers to turn promising ideas into shipped, maintained performance improvements.
- Own the process from profiling ambiguous bottlenecks through kernel or integration changes, benchmarked results, and ongoing maintenance.
Requirements
- Authored custom CUDA kernels beyond merely calling cuDNN or cuBLAS.
- Strong understanding of GPU architecture and performance, including memory hierarchy, warps, shared memory, register pressure, bandwidth versus compute limits, occupancy, and tensor core utilization.
- Proficiency with low-level profiling using Nsight Systems and/or Nsight Compute.
- Strong C/C++ skills.
- CUTLASS experience and tensor core utilization strategies are preferred.
- Triton kernel experience and/or PyTorch custom operation integration is preferred.
- Experience building benchmark harnesses and performance regression tests is preferred.
Benefits
- Open to locations beyond preferred San Francisco and Boston.
- 100% of medical, dental, and vision premiums covered for employees and dependents.
- 401(k) matching up to 4% of base pay.
- Unlimited PTO and company-wide Refill Days.
- Competitive base salary with equity.
About Liquid AI
We build efficient general-purpose AI at every scale.
