about 3 hours ago
Responsibilities
- Design and operate distributed training infrastructure for neural operator architectures.
- Optimize training pipelines for throughput, fault tolerance, and cost efficiency.
- Build and maintain experiment tracking and observability systems.
- Solve data loading bottlenecks for large-scale mesh datasets.
- Optimize data pipelines for efficient I/O from cloud storage.
- Build serving infrastructure for pre-trained models.
- Design and implement model packaging pipelines for customer deployment.
- Improve developer experience for the Research team with reliable CI/CD.
Requirements
- 5+ years of experience building and operating ML infrastructure at scale.
- Deep expertise in distributed training and debugging NCCL hangs.
- Strong systems fundamentals including Linux, networking, and storage I/O.
- Production experience with Kubernetes and SLURM for job orchestration.
- Proficiency in Python and ML frameworks, preferably PyTorch.
- Experience with cloud GPU infrastructure, ideally CoreWeave.
Benefits
- Equity options to share in the company's success.
- 10% employer pension contribution for future investment.
- Free office lunches to keep you energized.
- Enhanced parental leave with full pay for maternity and paternity.
- 25 days of annual leave plus public holidays.
- Private medical insurance with 100% employee cover.
- Personal development support for learning and growth.
- Employee Assistance Programme for confidential wellbeing support.
