1 year ago
Responsibilities
- Design and write high-performing, scalable software for model training.
- Develop tools that support and accelerate research and LLM training.
- Coordinate with infrastructure, efficiency, serving, and scientific teams to build an integrated post-training ecosystem.
- Implement techniques to improve performance and accelerate SFT, offline preference, and RL training cycles.
- Research, implement, and experiment with ideas on cluster and data infrastructure.
- Contribute production code and support team research efforts.
Requirements
- Extremely strong software engineering skills.
- Proficiency in Python and related ML frameworks such as JAX, PyTorch, and/or XLA/MLIR.
- Experience using and debugging large-scale distributed training strategies, including memory and speed profiling.
- Experience with distributed training infrastructure such as Kubernetes and associated frameworks such as Ray is a bonus.
- Hands-on experience with model post-training, with an emphasis on scalability and performance, is a bonus.
- Experience in ML, LLM, and RL academic research is a bonus.
- Ability to work with complex ML codebases and collaborate with scientists, engineers, and colleagues with varying levels of software engineering experience.
Benefits
- Open and inclusive culture and work environment
- Work with a team on cutting-edge AI research
- Weekly lunch stipend, in-office lunches, and snacks
- Full health and dental benefits, including a separate mental health budget
- 100% parental leave top-up for up to six months
- Personal enrichment benefits for arts and culture, fitness and well-being, quality time, and workspace improvement
- Remote-flexible work with offices in Toronto, New York, San Francisco, London, and Paris
- Co-working stipend
- Six weeks of vacation, equivalent to 30 working days
Tech Stack
Categories
AI ResearchML Engineering
About Cohere
Cohere builds large language models and an enterprise AI platform that companies use for search, summarization, and workflow automation, delivered via API or private deployments. Founded in 2019 and headquartered in Toronto, it focuses on multilingual models, data controls, and options to run across major clouds or on-premises. The business is privately held and serves security- and compliance-sensitive organizations.
