26 days ago
Los Altos, CA, USASenior
Base Salary
$120k - $222k/yr
Responsibilities
- Develop MAX’s API surface, including namespaces, abstractions, extension points, programming models, compatibility guarantees, and developer experience.
- Design inference APIs supporting model loading, distributed inference, quantization, tokenization, pipelines, batching, streaming, and deployment ergonomics.
- Develop training APIs for single-device and distributed execution, including device placement, parallelism, checkpointing, and observability.
- Align programming conventions, naming, and types across Python and Mojo-adjacent surfaces.
- Write, socialize, and iterate on technical RFCs and specifications with internal engineers and external users.
- Partner with compiler/runtime, kernels, cloud/serving, and documentation/DevRel teams.
- Create reference implementations, examples, architecture templates, and best-practice patterns.
Requirements
- Experience or strong interest in designing developer-facing APIs for SDKs, frameworks, or platforms.
- Good understanding of modern AI framework design tradeoffs, including PyTorch, JAX, TensorFlow, vLLM, and XLA/MLIR-adjacent ecosystems.
- Proficiency in one or more systems or performance languages such as C++, Rust, or Go.
- Proficiency in Python; familiarity with Mojo is a plus.
- Strong API design instincts involving naming, composability, types, error handling, configurability, and clarity.
- Strong written communication skills for producing implementable technical specifications.
- Commitment to pragmatic engineering standards and incremental development.
Benefits
- Comprehensive healthcare coverage may be available, depending on location.
- Retirement and savings programs, paid time off, wellbeing resources, family support programs, and learning and development opportunities may be included.
- Employee stock purchase opportunities may be available.
- Los Altos, California headquarters location with a minimum of three days per week onsite.
- Relocation assistance is provided for out-of-state US candidates.
- Regular team onsites and local meetups are organized, with expected travel two to four times per year.
About Modular
Modular builds an AI developer platform for training and especially inference/serving, centered on the MAX runtime and the Mojo programming language. Its tools accelerate and deploy models from frameworks like PyTorch and TensorFlow on CPUs and GPUs, for teams running on cloud or on‑prem infrastructure. Founded in 2022, the company operates remote‑first with an office in Los Altos, CA, and sells a commercial platform and enterprise support to organizations productionizing generative and classical ML.