Modular

Software Engineer, Hardware Enablement

Modular
Apply
6 months ago
Remote, United States +2 moreSenior

Base Salary

$180k - $270k/yr

Responsibilities

  • Bring up and validate support for new hardware architectures across kernels, compiler infrastructure, runtime, graph execution, and model serving.
  • Write, port, and optimize Mojo kernels for novel accelerator architectures.
  • Investigate correctness and performance issues involving kernel execution, memory movement, compiler lowering, graph decisions, synchronization, and vendor runtime behavior.
  • Analyze accelerator ISAs, execution models, memory hierarchies, synchronization mechanisms, compiler constraints, and vendor toolchains.
  • Map AI operators and workloads such as matrix multiplication, convolution, and reductions onto target architectures.
  • Develop portability infrastructure, compiler integrations, tooling, testing, debugging workflows, and integration tests for additional hardware platforms.
  • Collaborate with hardware vendors and silicon engineering teams on platform capabilities and software integration.
  • Benchmark and profile workloads to identify bottlenecks and improve production performance.
  • Document emerging architectures through technical documentation, demos, design discussions, and engineering write-ups.
  • Participate in company events, onsites, hackathons, and team collaboration activities.

Requirements

  • 5+ years of experience in high-performance computing, compiler engineering, accelerator software, systems engineering, or a closely related industry or research domain.
  • Strong C++ skills and experience contributing to complex, multi-component software systems.
  • Hands-on experience with at least one heterogeneous programming model such as CUDA, SYCL, OpenCL, or a comparable accelerator programming environment.
  • Experience writing or modifying GPU kernels, custom operators, accelerator code, or working with PyTorch at the C++ or systems layer.
  • Working knowledge of memory hierarchy, parallel execution, synchronization, data movement, and software performance relationships to hardware architecture.
  • Experience debugging issues across software abstraction boundaries and systematically reasoning about correctness and performance.
  • Ability to learn unfamiliar hardware platforms and read architecture manuals, ISA documentation, and vendor technical materials.
  • Strong communication and collaboration skills across compiler, runtime, kernel, and hardware teams.
  • Experience with non-GPU accelerators such as DSPs, NPUs, or AI ASICs is helpful but not required.
  • Familiarity with MLIR, LLVM, Triton, CUTLASS, CuTe, graph compilers, model serving, inference optimization, runtime concepts, portability layers, or hardware-specific performance tools is helpful but not required.

Benefits

  • Candidates may work from home remotely or from offices in Los Altos, CA, or Edinburgh; candidates in the US, Canada, and UK are welcome.
  • New-hire onboarding is conducted in person at the appropriate office.
  • Benefits may include healthcare coverage, retirement and savings programs, employee stock purchase opportunities, paid time off, wellbeing resources, family support, and learning and development opportunities.
  • The company organizes regular onsites and local meetups, and travel 2–4 times per year is expected.
  • Compensation may also include RSU grants, annual target bonus, equity, and other benefits.

Tech Stack

Categories

Modular

About Modular

201-500 employees

Modular builds an AI developer platform for training and especially inference/serving, centered on the MAX runtime and the Mojo programming language. Its tools accelerate and deploy models from frameworks like PyTorch and TensorFlow on CPUs and GPUs, for teams running on cloud or on‑prem infrastructure. Founded in 2022, the company operates remote‑first with an office in Los Altos, CA, and sells a commercial platform and enterprise support to organizations productionizing generative and classical ML.

Contact me