10 hours ago
Base Salary
$250k - $300k/yr
Responsibilities
- Automate integration of CPU, GPU and storage capacity from diverse hardware providers.
- Build control-plane automation for provisioning, imaging, monitoring, repairing and recovering machines.
- Detect and remediate unhealthy GPUs, thermals, disks and other hardware failures.
- Manage machine images, network boot, kernels, firmware and bare-metal network configuration.
- Develop fleet health monitoring, hardware acceptance testing and benchmarking across CPU, disk, GPU, interconnect and network systems.
- Debug issues across hardware, operating-system, networking, container-runtime and Python control-plane layers.
- Participate in the on-call rotation and respond to production incidents.
Requirements
- 5+ years of experience writing high-quality production code.
- Experience operating physical hardware fleets, including bare-metal provisioning, BMC/IPMI, PXE or network boot and firmware, or building control planes that manage them.
- Strong cloud skills.
- Strong knowledge of low-level operating-system foundations, including the Linux kernel, drivers, networking, file systems and containers.
- Ability to debug across layers, including BGP flapping, Linux RPS, vBIOS bugs and Python control-plane services.
- Willingness to participate in on-call rotations and respond to production incidents.
- Experience with GPUs and the NVIDIA software stack in production, including drivers, health monitoring, XIDs, RDMA or NVLink, is preferred.
- Prior experience with Go is preferred.
About Modal
Modal builds a serverless compute platform for AI and data workloads, offering instant GPU access, sub-second container starts, and native storage to run inference, fine-tuning, and batch jobs. It sells a usage-based cloud service to developers and ML teams to deploy generative models and pipelines. Privately held and headquartered in New York City, its customers include companies like DoorDash and Ramp.
