10 months ago
Responsibilities
- Build and own the training framework for large-scale LLM training.
- Design distributed training abstractions for data, tensor, and pipeline parallelism, including FSDP/ZeRO strategies, memory management, and checkpointing.
- Improve training throughput and stability on multi-node GB200/300, AMD, H200/100, and similar clusters.
- Develop and maintain tooling for monitoring, logging, debugging, and developer ergonomics.
- Collaborate with infrastructure teams on cluster, container, and hardware configurations for high-performance training.
- Investigate and resolve performance bottlenecks across the ML systems stack.
- Build robust systems for reproducible, debuggable, large-scale training runs.
- Develop data-loading and caching pipelines, performance profiling, internal metrics and monitoring, regression-testing infrastructure, and fault-tolerant distributed checkpointing systems.
Requirements
- Strong engineering experience in large-scale distributed training or HPC systems.
- Deep familiarity with JAX internals, distributed training libraries, or custom kernels and fused operations.
- Experience with multi-node cluster orchestration using Slurm, Ray, Kubernetes, or similar tools.
- Ability to debug performance issues across CUDA, NCCL, networking, I/O, and data pipelines.
- Experience with containerized environments such as Docker and Singularity/Apptainer.
- Track record of building tools that improve developer velocity for ML teams.
- Experience training LLMs or other large transformer architectures is preferred.
- Contributions to ML frameworks such as PyTorch, JAX, DeepSpeed, Megatron, or xFormers are preferred.
- Familiarity with evaluation and serving frameworks such as vLLM and TensorRT-LLM, including custom KV caches, is preferred.
- Experience with data pipeline optimization, sharded datasets, or caching strategies is preferred.
- Background in performance engineering, profiling, or low-level systems is preferred.
- A paper at a top-tier venue such as NeurIPS, ICML, ICLR, AIStats, MLSys, JMLR, AAAI, Nature, COLING, ACL, or EMNLP is a bonus.
Benefits
- Weekly lunch stipend, in-office lunches and snacks.
- Full health and dental benefits, including a separate mental-health budget.
- 100% parental-leave top-up for up to six months.
- Personal enrichment benefits for arts and culture, fitness and well-being, quality time, and workspace improvement.
- Remote-flexible work with offices in Toronto, New York, San Francisco, London, and Paris, plus a coworking stipend.
- Six weeks of vacation, totaling 30 working days.
- Inclusive work environment and accommodations support during the recruitment process.
Tech Stack
Categories
About Cohere
Cohere builds large language models and an enterprise AI platform that companies use for search, summarization, and workflow automation, delivered via API or private deployments. Founded in 2019 and headquartered in Toronto, it focuses on multilingual models, data controls, and options to run across major clouds or on-premises. The business is privately held and serves security- and compliance-sensitive organizations.
