Modular

Inference Optimization Engineer

Modular
Apply
3 months ago
Remote, United StatesSenior

Base Salary

$198k - $286k/yr

Responsibilities

  • Build the optimization platform that improves LLM inference performance across GPU and ASIC architectures.
  • Profile customer inference workloads and apply optimizations across kernels, inference engines, distributed systems, and cloud infrastructure.
  • Develop reusable tooling and libraries that turn performance improvements into an automated optimization loop.
  • Define technical direction for Modular Cloud and maintain LLM performance at the Pareto frontier for agentic use cases.
  • Partner with GTM on customized customer inference solutions and collaborate with engineering and product teams to deliver optimizations in production.
  • Publish technical blog posts on LLM inference optimization approaches and industry best practices.

Requirements

  • 5+ years of experience in distributed systems or performance engineering.
  • Demonstrated experience building durable, reusable software tools and libraries adopted across teams and functions.
  • Strong technical judgment, prioritization, communication, and technical leadership skills.
  • Creativity, curiosity, collaboration, and alignment with the company culture.

Benefits

  • Comprehensive healthcare coverage, retirement and savings programs, employee stock purchase opportunities, paid time off, wellbeing resources, family support programs, and learning and development opportunities may be available, varying by location.
  • Competitive compensation may include annual target bonus, equity, RSU grants, and benefits.
  • Remote work from home or office work in Los Altos, California is available for candidates in the US or Canada.
  • New-hire onboarding is conducted in person at the Los Altos, California office.
  • Regular team onsites and local meetups are provided, with travel expected 2–4 times per year.
  • Specific benefits vary by location.
Modular

About Modular

201-500 employees

Modular builds an AI developer platform for training and especially inference/serving, centered on the MAX runtime and the Mojo programming language. Its tools accelerate and deploy models from frameworks like PyTorch and TensorFlow on CPUs and GPUs, for teams running on cloud or on‑prem infrastructure. Founded in 2022, the company operates remote‑first with an office in Los Altos, CA, and sells a commercial platform and enterprise support to organizations productionizing generative and classical ML.

Contact me