Modular

Senior AI Kernel Engineer

Modular
Apply
8 months ago
Remote, United StatesSenior

Base Salary

$180k - $270k/yr

Responsibilities

  • Design, implement, and optimize performance-critical kernels for AI inference workloads, including GEMM, attention, communication, and fusion.
  • Lead kernel optimization across single-GPU, multi-GPU, and heterogeneous hardware environments.
  • Balance latency, throughput, memory footprint, and numerical precision in kernel implementations.
  • Drive adoption of hardware features such as Tensor Cores, asynchronous execution, and advanced memory spaces.
  • Use profilers, hardware counters, and microbenchmarks to identify and deliver performance improvements.
  • Collaborate with compiler and runtime teams on code generation, scheduling, and kernel fusion strategies.
  • Review and mentor engineers on kernel design and performance tuning.
  • Contribute to technical roadmaps and long-term AI inference performance strategy.

Requirements

  • 5+ years of experience in performance-critical systems or kernel development, or equivalent depth of expertise.
  • Strong proficiency in C/C++ and low-level programming.
  • Extensive hands-on experience with GPU kernel programming using CUDA, HIP, or equivalent.
  • Deep understanding of GPU architecture, including memory hierarchies, synchronization, and execution models.
  • Proven record of delivering measurable performance improvements in production systems.
  • Strong problem-solving skills and ability to work independently on complex performance challenges.
  • Helpful qualifications include PTX, assembly-level tuning, code generation frameworks such as Triton, distributed or multi-GPU inference, custom AI accelerators, transformers, LLMs, diffusion models, open-source kernel or compiler contributions, and collaboration with hardware or compiler teams.

Benefits

  • Candidates may work from home remotely in the US or Canada or from the Los Altos, California office.
  • New-hire onboarding is conducted in person at the Los Altos, California office.
  • Benefits may include healthcare coverage, retirement and savings programs, employee stock purchase opportunities, paid time off, wellbeing resources, family support programs, and learning and development opportunities.
  • Compensation may include RSU grants, annual target bonus, equity, and benefits.
  • Regular team onsites and local meetups are provided, with expected travel 2-4 times per year.
  • Specific benefit packages may vary by location.

Tech Stack

Categories

Modular

About Modular

201-500 employees

Modular builds an AI developer platform for training and especially inference/serving, centered on the MAX runtime and the Mojo programming language. Its tools accelerate and deploy models from frameworks like PyTorch and TensorFlow on CPUs and GPUs, for teams running on cloud or on‑prem infrastructure. Founded in 2022, the company operates remote‑first with an office in Los Altos, CA, and sells a commercial platform and enterprise support to organizations productionizing generative and classical ML.

Contact me