Software Engineer – AI Inference Engine
FriendliAI6 months ago
Responsibilities
- Design and optimize custom GPU kernels for transformer and diffusion AI workloads.
- Develop core inference-engine components including the kernel compiler, memory planner, runtime, and related infrastructure.
- Collaborate with cloud and infrastructure engineers to improve end-to-end inference performance.
- Analyze software and hardware performance bottlenecks and implement targeted optimizations.
- Add support for new model architectures and tensor compute patterns.
- Maintain profiling, benchmarking, and validation tools for production-grade performance infrastructure.
Requirements
- 5+ years of experience in production or high-impact research environments.
- Production-level expertise in Python and C++.
- Bachelor’s or master’s degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent.
- Experience developing machine learning frameworks or performance-critical runtime systems.
- Hands-on experience writing, optimizing, and profiling GPU kernels.
- Experience working with generative AI models such as transformer and diffusion models.
- Preferred: experience with machine learning compilers or code generation systems, dynamic shape compilation, memory planning, kernel fusion, inference engines, compilers, high-performance numerical libraries, and multi-GPU or distributed inference strategies.
Benefits
- Flexible working hours.
- Daily lunch and dinner, plus unlimited snacks and beverages.
- Supportive and highly collaborative work environment.
- Health check-up support and top-tier equipment and hardware support.
- Startup equity, health insurance, competitive compensation, and other benefits.
- Opportunity to work on generative AI infrastructure at a fast-moving AI company.