Modular

AI Runtime Engineer

Modular
Apply
5 hours ago
Edinburgh, United KingdomMid Level

Responsibilities

  • Design and develop runtime and cross-stack optimizations for CPU, GPU, and accelerator efficiency.
  • Optimize CPU overhead, caching, data locality, and data transfer across multiple hardware topologies.
  • Collaborate with compiler, kernels, serving, models, tooling, infrastructure, and customer success teams.
  • Engage with customers to understand performance requirements and use cases.
  • Design automated performance analysis and benchmarking systems.

Requirements

  • 2+ years of experience working on high-performance computing systems.
  • Experience programming in C++ and working with complex software systems.
  • Experience with CPU or GPU runtime optimization and performance analysis on CPUs, GPUs, or AI accelerators.
  • Proficiency with one or more CPU or GPU profiling tools.
  • Helpful qualifications include ML graph optimization, parallel or distributed programming, heterogeneous ML computation, code generation, and exposure to MLIR, LLVM, or Mojo.
  • An advanced degree in Computer Science or a related area is a plus.

Benefits

  • Healthcare coverage, retirement and savings programs, employee stock purchase opportunities, paid time off, wellbeing resources, family support programs, and learning and development opportunities may be included depending on location.
  • RSU grants, annual target bonus, equity, and other benefits are part of the total compensation package.
  • The role is based in the Edinburgh office with a minimum of 3 days per week on-site.
  • New hires complete onboarding in person, and relocation assistance is available for eligible candidates.
  • Regular team onsites and local meetups are offered, with travel 2–4 times per year expected.
  • Candidates based in the United Kingdom are welcome to apply.

Tech Stack

Modular

About Modular

201-500 employees

Modular builds an AI developer platform for training and especially inference/serving, centered on the MAX runtime and the Mojo programming language. Its tools accelerate and deploy models from frameworks like PyTorch and TensorFlow on CPUs and GPUs, for teams running on cloud or on‑prem infrastructure. Founded in 2022, the company operates remote‑first with an office in Los Altos, CA, and sells a commercial platform and enterprise support to organizations productionizing generative and classical ML.

Contact me