9 months ago
Responsibilities
- Build and own the training framework for large-scale LLM training.
- Design distributed training abstractions for data, tensor, and pipeline parallelism, including FSDP/ZeRO strategies, memory management, and checkpointing.
- Improve training throughput and stability on multi-node GB200/300, AMD, H200/100, and similar clusters.
- Develop and maintain tooling for monitoring, logging, debugging, and developer ergonomics.
- Collaborate with infrastructure teams on cluster, container, and hardware configurations for high-performance training.
- Investigate and resolve performance bottlenecks across the ML systems stack.
- Build robust systems for reproducible, debuggable, large-scale training runs.
- Develop data-loading and caching pipelines, performance profiling, internal metrics and monitoring, regression-testing infrastructure, and fault-tolerant distributed checkpointing systems.
Requirements
- Strong engineering experience in large-scale distributed training or HPC systems.
- Deep familiarity with JAX internals, distributed training libraries, or custom kernels and fused operations.
- Experience with multi-node cluster orchestration using Slurm, Ray, Kubernetes, or similar tools.
- Ability to debug performance issues across CUDA, NCCL, networking, I/O, and data pipelines.
- Experience with containerized environments such as Docker and Singularity/Apptainer.
- Track record of building tools that improve developer velocity for ML teams.
- Experience training LLMs or other large transformer architectures is preferred.
- Contributions to ML frameworks such as PyTorch, JAX, DeepSpeed, Megatron, or xFormers are preferred.
- Familiarity with evaluation and serving frameworks such as vLLM and TensorRT-LLM, including custom KV caches, is preferred.
- Experience with data pipeline optimization, sharded datasets, or caching strategies is preferred.
- Background in performance engineering, profiling, or low-level systems is preferred.
- A paper at a top-tier venue such as NeurIPS, ICML, ICLR, AIStats, MLSys, JMLR, AAAI, Nature, COLING, ACL, or EMNLP is a bonus.
Benefits
- Weekly lunch stipend, in-office lunches and snacks.
- Full health and dental benefits, including a separate mental-health budget.
- 100% parental-leave top-up for up to six months.
- Personal enrichment benefits for arts and culture, fitness and well-being, quality time, and workspace improvement.
- Remote-flexible work with offices in Toronto, New York, San Francisco, London, and Paris, plus a coworking stipend.
- Six weeks of vacation, totaling 30 working days.
- Inclusive work environment and accommodations support during the recruitment process.
Tech Stack
Categories
About Cohere
Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation models and end-to-end AI products designed to solve real-world business problems. We partner closely with companies to deliver seamless integration, full customization, and easy-to-use solutions for their workforce and customers. Our all-in-one platform offers enterprises the highest levels of data security, privacy and optionality to deploy across all major cloud providers, private cloud environments, or on-premises. HQ: 171 John Street, 2nd Floor, Toronto, ON M5T 1X3
