over 1 year ago
Boston, MA, USASenior
Responsibilities
- Design, implement, and maintain large-scale model training infrastructure for job orchestration, scheduling, checkpointing, and experiment tracking.
- Build intuitive tooling for launching, monitoring, debugging, and reproducing research experiments.
- Scale reinforcement learning and machine learning pipelines across distributed compute clusters.
- Manage the efficient allocation and utilization of cloud-based compute resources.
- Partner with researchers to develop scalable training infrastructure, guide training-at-scale best practices, and contribute to core JAX model and training code.
- Build automated testing pipelines, CI/CD workflows for machine learning, and custom logging and telemetry systems.
Requirements
- Bachelor’s degree or higher in Computer Science, Computer Engineering, Machine Learning, or a related technical field.
- Strong software engineering fundamentals and a proven track record building ML training infrastructure, internal developer platforms, or scalable systems.
- Hands-on experience with large-scale training using JAX, PyTorch, or TensorFlow.
- Familiarity with distributed training, multi-host systems, data pipelines, and workloads on cloud platforms or orchestration systems such as Kubernetes, SLURM, GCP, or AWS.
- Strong cross-functional communication, ownership, and developer-experience orientation.
- Preferred experience in robotics, reinforcement learning, or other machine learning systems.
- Preferred experience designing abstractions that balance researcher flexibility with system reliability.
