1 year ago
Remote, United States or Remote, CanadaMid Level
Base Salary
$148k - $270k/yr
Responsibilities
- Implement and validate support for new hardware architectures across the Modular software stack.
- Write and optimize Mojo kernels for novel accelerator architectures.
- Improve portability infrastructure, tooling, and debugging workflows for target hardware.
- Collaborate with hardware vendors on platform understanding, integration tests, and issue triage.
- Study accelerator ISAs, memory hierarchies, and vendor toolchains, and share findings through demos and write-ups.
Requirements
- At least 2 years of experience in high-performance computing, compiler engineering, or a related industry or research domain.
- Proficiency in C++ and experience with complex, multi-component software systems.
- Hands-on experience with a heterogeneous programming model such as CUDA, SYCL, or OpenCL.
- Ability to learn new hardware platforms and read architecture manuals and vendor documentation.
- Experience with non-GPU accelerators, GPU kernels or custom operators, PyTorch at the C++ layer, Triton, CUTLASS, CuTe, MLIR, LLVM, hardware vendor teams, platform bring-up, model serving, or inference optimization is preferred but not required.
Benefits
- Healthcare coverage, retirement and savings programs, employee stock purchase opportunities, paid time off, wellbeing resources, family support programs, and learning and development opportunities may be provided depending on location.
- The role offers relocation assistance for eligible US candidates based outside Los Altos, CA.
- Earlier-career employees work hybrid from the Los Altos, CA office at least 3 days per week; senior members may work in-office or remotely.
- Onboarding is conducted in Los Altos, CA, travel is expected 2–4 times per year, and the company organizes onsites, local meetups, and hackathons.
About Modular
Modular builds an AI developer platform for training and especially inference/serving, centered on the MAX runtime and the Mojo programming language. Its tools accelerate and deploy models from frameworks like PyTorch and TensorFlow on CPUs and GPUs, for teams running on cloud or on‑prem infrastructure. Founded in 2022, the company operates remote‑first with an office in Los Altos, CA, and sells a commercial platform and enterprise support to organizations productionizing generative and classical ML.