FriendliAI

Software Engineer – AI Inference Engine

FriendliAI
Apply
6 months ago

Responsibilities

  • Design and optimize custom GPU kernels for transformer and diffusion AI workloads.
  • Develop core inference-engine components including the kernel compiler, memory planner, runtime, and related infrastructure.
  • Collaborate with cloud and infrastructure engineers to improve end-to-end inference performance.
  • Analyze software and hardware performance bottlenecks and implement targeted optimizations.
  • Add support for new model architectures and tensor compute patterns.
  • Maintain profiling, benchmarking, and validation tools for production-grade performance infrastructure.

Requirements

  • 5+ years of experience in production or high-impact research environments.
  • Production-level expertise in Python and C++.
  • Bachelor’s or master’s degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent.
  • Experience developing machine learning frameworks or performance-critical runtime systems.
  • Hands-on experience writing, optimizing, and profiling GPU kernels.
  • Experience working with generative AI models such as transformer and diffusion models.
  • Preferred: experience with machine learning compilers or code generation systems, dynamic shape compilation, memory planning, kernel fusion, inference engines, compilers, high-performance numerical libraries, and multi-GPU or distributed inference strategies.

Benefits

  • Flexible working hours.
  • Daily lunch and dinner, plus unlimited snacks and beverages.
  • Supportive and highly collaborative work environment.
  • Health check-up support and top-tier equipment and hardware support.
  • Startup equity, health insurance, competitive compensation, and other benefits.
  • Opportunity to work on generative AI infrastructure at a fast-moving AI company.

Tech Stack

Categories

FriendliAI

About FriendliAI

51-200 employees
Contact me