Modal

Member of Technical Staff - ML Performance

Modal
Apply
4 months ago
San Francisco, CA, USA or New York, NY, USASenior
H1B Sponsor

Base Salary

$150k - $350k/yr

Responsibilities

  • Contribute to open-source projects and Modal’s container runtime.
  • Improve throughput and reduce latency for language and diffusion models.
  • Optimize machine learning systems and GPU performance at scale.

Requirements

  • At least five years of experience writing high-quality, high-performance code.
  • Experience with torch, high-level machine learning frameworks, and inference engines such as vLLM or TensorRT.
  • Familiarity with NVIDIA GPU architecture and CUDA.
  • Experience with machine learning performance engineering, including debugging SM occupancy issues, making algorithms compute-bound, and eliminating host overhead.
  • Familiarity with low-level operating-system foundations such as the Linux kernel, file systems, and containers is a nice-to-have.

Tech Stack

LinuxSeaborn

Categories

Modal

About Modal

51-200 employees

Customers rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. Every era of computing came with new workloads that previous infrastructure couldn't serve: mainframes, databases, the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice. The window to build is open right now.