over 1 year ago
Responsibilities
- Design and write high-performance, scalable software for model training
- Analyze how architectural modifications and design choices affect training throughput and quality
- Write low-level CUDA and Triton kernels to optimize accelerator performance
- Research, implement, and experiment with ideas across supercompute and data infrastructure
- Develop performance and profiling tools to identify and remove bottlenecks
- Collaborate with researchers on efficient and reliable language understanding and generation systems
Requirements
- Extremely strong software engineering skills
- Proficiency in Python and related machine-learning frameworks including JAX, PyTorch, and XLA/MLIR
- Experience writing GPU kernels using CUDA, Triton, or similar technologies
- Experience with large-scale distributed training strategies
- Familiarity with autoregressive sequence models such as Transformers
- A paper at a top-tier venue such as NeurIPS, ICML, ICLR, AIStats, MLSys, JMLR, AAAI, Nature, COLING, ACL, or EMNLP is a bonus
Benefits
- Remote-flexible work with offices in Toronto, New York, San Francisco, London, and Paris, plus a co-working stipend
- Weekly lunch stipend, in-office lunches, and snacks
- Full health and dental benefits, including a separate mental health budget
- 100% parental leave top-up for up to 6 months
- Personal enrichment benefits for arts and culture, fitness and well-being, quality time, and workspace improvement
- Six weeks of vacation, or 30 working days
- Inclusive work environment and reasonable accommodations during recruitment
Categories
About Cohere
Cohere builds large language models and an enterprise AI platform that companies use for search, summarization, and workflow automation, delivered via API or private deployments. Founded in 2019 and headquartered in Toronto, it focuses on multilingual models, data controls, and options to run across major clouds or on-premises. The business is privately held and serves security- and compliance-sensitive organizations.
