6 months ago
Base Salary
$150k - $350k/yr
Responsibilities
- Contribute to open-source projects and Modal’s container runtime.
- Improve throughput and reduce latency for language and diffusion models.
- Optimize machine learning systems and GPU performance at scale.
Requirements
- At least five years of experience writing high-quality, high-performance code.
- Experience with torch, high-level machine learning frameworks, and inference engines such as vLLM or TensorRT.
- Familiarity with NVIDIA GPU architecture and CUDA.
- Experience with machine learning performance engineering, including debugging SM occupancy issues, making algorithms compute-bound, and eliminating host overhead.
- Familiarity with low-level operating-system foundations such as the Linux kernel, file systems, and containers is a nice-to-have.
Tech Stack
LinuxSeaborn
Categories
About Modal
Modal builds a serverless compute platform for AI and data workloads, offering instant GPU access, sub-second container starts, and native storage to run inference, fine-tuning, and batch jobs. It sells a usage-based cloud service to developers and ML teams to deploy generative models and pipelines. Privately held and headquartered in New York City, its customers include companies like DoorDash and Ramp.
