4 months ago
Base Salary
$150k - $350k/yr
Responsibilities
- Contribute to open-source projects and Modal’s container runtime.
- Improve throughput and reduce latency for language and diffusion models.
- Optimize machine learning systems and GPU performance at scale.
Requirements
- At least five years of experience writing high-quality, high-performance code.
- Experience with torch, high-level machine learning frameworks, and inference engines such as vLLM or TensorRT.
- Familiarity with NVIDIA GPU architecture and CUDA.
- Experience with machine learning performance engineering, including debugging SM occupancy issues, making algorithms compute-bound, and eliminating host overhead.
- Familiarity with low-level operating-system foundations such as the Linux kernel, file systems, and containers is a nice-to-have.
Tech Stack
LinuxSeaborn
Categories
About Modal
Customers rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. Every era of computing came with new workloads that previous infrastructure couldn't serve: mainframes, databases, the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice. The window to build is open right now.
