Modular

Senior AI Runtime Engineer

Modular
Apply
4 hours ago
Remote, United States or Remote, CanadaSenior

Base Salary

$216k - $324k/yr

Responsibilities

  • Design and develop runtime and cross-stack optimizations to improve CPU and GPU efficiency, including CPU overhead, caching, and data locality improvements.
  • Port the Modular runtime stack to new hardware platforms and develop an API to streamline the porting process.
  • Collaborate with compiler, kernels, serving, and models teams to build high-performance technologies for CPU and GPU hardware.
  • Engage with customers and the customer success team to understand performance requirements and use cases.
  • Collaborate with tooling and infrastructure teams to design automated performance analysis and benchmarking systems.

Requirements

  • 5+ years of experience working on high-performance computing systems.
  • Experience programming in C++ and working with complex software systems.
  • Experience with CPU or GPU runtime optimization and performance analysis on CPUs, GPUs, or AI accelerators.
  • Proficiency with one or more CPU or GPU profiling tools.
  • Helpful qualifications include experience with ML graph optimization, parallel or distributed programming, heterogeneous ML computation, code generation, MLIR, LLVM, or the Mojo programming language.
  • An advanced degree in Computer Science or a related field is helpful but not required.

Benefits

  • Benefits may include comprehensive healthcare coverage, retirement and savings programs, employee stock purchase opportunities, paid time off, wellbeing resources, family support programs, and learning and development opportunities.
  • The compensation package may include RSU grants, annual target bonus, equity, and benefits.
  • The role may be performed remotely from home in the US or Canada or from the Los Altos, California office.
  • New-hire onboarding is conducted in person at the Los Altos, California office.
  • Regular team onsites and local meetups are held, with expected travel two to four times per year.

Tech Stack

Categories

Modular

About Modular

201-500 employees

Modular builds an AI developer platform for training and especially inference/serving, centered on the MAX runtime and the Mojo programming language. Its tools accelerate and deploy models from frameworks like PyTorch and TensorFlow on CPUs and GPUs, for teams running on cloud or on‑prem infrastructure. Founded in 2022, the company operates remote‑first with an office in Los Altos, CA, and sells a commercial platform and enterprise support to organizations productionizing generative and classical ML.

Contact me