7 months ago
San Francisco, CA, USA or New York, NY, USASenior
Responsibilities
- Collaborate with ML researchers to improve training throughput and reliability.
- Plan and build GPU infrastructure with OEMs, cloud service providers, and other partners.
- Improve the density and scalability of compute environments for increasingly large reinforcement-learning workloads.
- Create software and systems to automate building, monitoring, and running GPU clusters.
- Build workload scheduling and data movement systems for Cursor’s training infrastructure.
Requirements
- Strong background in systems and infrastructure-focused software engineering, particularly with Python, TypeScript, Rust, and Golang.
- Experience with distributed storage and networking infrastructure on Linux systems across cloud and bare-metal environments.
- Exposure to large-scale systems, ideally across thousands of nodes with significant resource footprints.
- Production use of infrastructure-as-code and configuration management across hosts and Kubernetes.
- Experience with Nvidia GPUs, InfiniBand or RoCE, Blackwell or Hopper hardware, Ray, or Slurm is a nice-to-have.
