6 hours ago
Base Salary
$220k - $300k/yr
Responsibilities
- Improve container startup, checkpoint, and restore performance for large inference and training workloads.
- Build multi-GPU and accelerator-aware snapshotting for GPU memory and RDMA-enabled workloads.
- Design zero-copy and low-copy data paths across container memory, filesystems, storage, and Modal’s runtime.
- Optimize snapshot pipelines using parallel uploads, direct I/O, incremental snapshots, and improved memory handling.
- Improve container image and filesystem performance across EROFS, FUSE, page caches, overlay filesystems, and remote storage.
- Extend the sandboxed runtime for new GPUs, drivers, profiling tools, and device capabilities across NVIDIA and AMD hardware.
- Debug failures involving system calls, virtual memory, process lifecycle, kernel behavior, GPU drivers, and container isolation.
- Roll out runtime and kernel changes using compatibility controls, scheduling constraints, feature flags, and observability.
- Work across runtime, scheduler, storage, and GPU infrastructure from investigation through production deployment.
- Contribute to gVisor and upstream runtime technology when appropriate.
Requirements
- Strong Linux systems knowledge covering processes, virtual memory, filesystems, system calls, scheduling, namespaces, cgroups, and signals.
- Experience building or debugging container runtimes, sandboxes, the Linux kernel, or similarly low-level infrastructure.
- Strong programming ability in Rust, Go, or another systems language, with interest in becoming productive in both Rust and Go.
- Experience profiling and improving systems where memory movement, I/O, synchronization, or kernel interactions dominate performance.
- Ability to debug across application, runtime, driver, and kernel abstraction boundaries.
- Ability to turn loosely defined production problems into reliable systems.
- Helpful but not required experience includes gVisor, runsc, runc, OCI runtimes, seccomp, checkpoint/restore systems, Linux kernel development, virtualization, sandboxing, kernel modules, device proxying, CUDA, ROCm, GPU drivers, accelerator virtualization, GPU profiling, RDMA, high-performance networking, FUSE, EROFS, direct I/O, mmap, page-cache behavior, storage engines, large-memory or multi-GPU inference and training systems, and open-source systems software.
Benefits
- Opportunity to shape the architecture of foundational container runtime infrastructure used by Modal’s inference and training workloads.
- Opportunity to contribute to open-source runtime technology and solve deep systems problems with direct production impact.
About Modal
Modal builds a serverless compute platform for AI and data workloads, offering instant GPU access, sub-second container starts, and native storage to run inference, fine-tuning, and batch jobs. It sells a usage-based cloud service to developers and ML teams to deploy generative models and pipelines. Privately held and headquartered in New York City, its customers include companies like DoorDash and Ramp.
