12 hours ago
Remote, United States +2 moreSenior
Base Salary
$180k - $270k/yr
Responsibilities
- Design and implement high-performance attention kernels for prefill and decode, including MHA, GQA, paged attention, sliding-window/global attention, and fused FlashAttention-style algorithms.
- Build and optimize general AI kernels such as matrix multiplication, softmax, normalization, activation, embedding, reduction, and data-movement primitives.
- Contribute to Modular’s Attention framework and develop reusable interfaces supporting portability, composability, tunability, and hardware-specific optimization.
- Profile workloads, identify bottlenecks, establish performance models, and tune implementations against hardware limits and competitive baselines.
- Develop benchmarks, correctness tests, regression tests, and automated performance analysis across model shapes and hardware generations.
- Collaborate with Modular teams and hardware partners on software-stack improvements, architecture validation, toolchains, low-level issues, and hardware/software interfaces.
- Write design documents, participate in technical reviews, share performance insights, and mentor teammates.
Requirements
- 5+ years of relevant industry or research experience in high-performance computing, AI kernel development, accelerator programming, compiler engineering, or a related field.
- Demonstrated experience writing and optimizing production-quality GPU or accelerator kernels.
- Strong understanding of parallel algorithms, numerical behavior, memory access patterns, synchronization, and data movement.
- Strong knowledge of modern attention algorithms, including prefill and decode performance tradeoffs.
- Understanding of architectures such as NPUs, DSPs, TPUs, and custom ASICs, including compute units, memory hierarchies, and execution models.
- Experience with at least one heterogeneous programming model or kernel ecosystem such as CUDA, Triton, SYCL, OpenCL, HIP, or a vendor accelerator SDK.
- Strong performance-analysis skills and proficiency with profiling, tracing, benchmarking, and debugging tools.
- Ability to read architecture manuals, compiler output, and low-level generated code and systematically investigate performance gaps.
- Preferred qualifications include Mojo, MLIR, LLVM, FlashAttention, CuTe/CUTLASS, Pallas, PyTorch, JAX, TensorFlow, vLLM, SGLang, TensorRT-LLM, GPT, Gemma, performance modeling, autotuning, numerical validation, and experience bringing up kernels on new hardware platforms.
- An advanced degree in Computer Science, Electrical Engineering, Computer Engineering, or a related field is helpful but not required.
Benefits
- Benefits may include healthcare coverage, retirement and savings programs, employee stock purchase opportunities, paid time off, wellbeing resources, family support programs, and learning and development opportunities.
- The role can be based in an office in Los Altos, California, or Edinburgh, or worked remotely from home; onboarding is conducted in person at the appropriate office.
- Regular team onsites and local meetups are provided, with expected travel 2–4 times per year.
- Competitive compensation may include RSU grants, annual target bonus, equity, and benefits.
Tech Stack
PyTorchTensorFlow
Categories
About Modular
Modular builds an AI developer platform for training and especially inference/serving, centered on the MAX runtime and the Mojo programming language. Its tools accelerate and deploy models from frameworks like PyTorch and TensorFlow on CPUs and GPUs, for teams running on cloud or on‑prem infrastructure. Founded in 2022, the company operates remote‑first with an office in Los Altos, CA, and sells a commercial platform and enterprise support to organizations productionizing generative and classical ML.