about 3 hours ago
Responsibilities
- Own the reliability and operation of GPU clusters for training and research.
- Debug issues across compute, storage, networking, schedulers, and distributed workloads.
- Improve CPU, GPU, and storage utilization through better tooling and automation.
- Onboard and migrate workloads across GPU providers and hardware platforms.
- Build monitoring, validation, and platform abstractions to reduce operational work for researchers.
- Contribute to the long-term architecture of Liquid AI’s training infrastructure and GPU platform.
Requirements
- Strong software engineering experience with production-quality infrastructure tooling.
- Deep knowledge of distributed systems, Linux, networking, and storage.
- Experience operating a shared compute cluster or distributed training platform.
- Track record of supporting production users and creating durable solutions.
- Technical depth to collaborate effectively with senior research and infrastructure engineers.
- Experience with SLURM, Kubernetes, Ray, Hadoop, or other distributed compute platforms is a plus.
- Experience supporting GPU, HPC, or large-scale AI training infrastructure is a plus.
- Familiarity with distributed storage, cluster schedulers, cloud providers, or infrastructure control planes is a plus.
Benefits
- High-impact ownership of infrastructure affecting AI model training efficiency.
- Competitive base salary with equity in a unicorn-stage company.
- 100% coverage of medical, dental, and vision premiums for employees and dependents.
- 401(k) matching up to 4% of base pay.
- Unlimited PTO plus company-wide Refill Days throughout the year.
