8 months ago
Remote, United StatesSenior
Base Salary
$167k - $273k/yr
Responsibilities
- Build and ship a distributed LLM inference platform with disaggregated inference, multi-node deployment, distributed KV-cache management, high-throughput batching, and high-performance networking.
- Improve operational excellence through request-to-kernel observability, multi-cloud deployments, autoscaling, and cold-start optimizations.
- Integrate kernel, serving, and cluster optimizations with the kernels and generative AI teams to achieve high application performance.
- Build Helm charts, Kubernetes operators, and reusable deployment tooling.
- Collaborate across teams to develop durable software tools and libraries used across the organization.
Requirements
- At least 5 years of backend engineering experience.
- Experience with Kubernetes and operating personally owned services.
- Ability to create durable, reusable software tools and libraries leveraged across teams and functions.
- Experience with machine learning technologies and use cases.
- Strong collaboration, creativity, curiosity, and alignment with company cultural values.
- Helpful qualifications include high-performance computing or networking experience, large-scale ML inference infrastructure experience, and familiarity with Golang.
Benefits
- Comprehensive healthcare coverage may be available, along with retirement and savings programs, employee stock purchase opportunities, paid time off, wellbeing resources, family support programs, and learning and development opportunities.
- Competitive compensation may include RSU grants, annual target bonus, equity, and benefits.
- Regular team-building events, team onsites, and local meetups are provided.
- Work may be performed remotely from home in the US or Canada or from the Los Altos, California office; new-hire onboarding is conducted in person in Los Altos.
- Travel to team events or meetups two to four times per year is expected.
Tech Stack
Categories
About Modular
Modular builds an AI developer platform for training and especially inference/serving, centered on the MAX runtime and the Mojo programming language. Its tools accelerate and deploy models from frameworks like PyTorch and TensorFlow on CPUs and GPUs, for teams running on cloud or on‑prem infrastructure. Founded in 2022, the company operates remote‑first with an office in Los Altos, CA, and sells a commercial platform and enterprise support to organizations productionizing generative and classical ML.