9 months ago
Responsibilities
- Improve LLM inference performance across the model execution stack.
- Identify bottlenecks and develop optimizations that reduce latency and increase throughput.
- Collaborate with modeling and systems teams to experiment, measure, and ship inference improvements.
- Build expertise in GPU/CUDA optimization, kernel-level improvements, and model execution strategies for MoE and large-scale architectures.
Requirements
- At least 5 years of experience writing high-performance, production-quality code.
- Strong programming skills in C++ or Python; Rust and Go are also welcome.
- Experience working with large language models and familiarity with the LLM inference ecosystem, including vLLM or SGLang.
- Ability to diagnose and resolve performance bottlenecks across the model execution stack.
- Experience with GPU programming, CUDA, or low-level systems optimization is a plus.
- Experience with transformer language modeling, including MoE, speculative decoding, or KV-cache optimizations, is a plus.
- Experience scaling performance-critical distributed systems for computation, search, or storage is a plus.
Benefits
- Open and inclusive culture and work environment
- Weekly lunch stipend, in-office lunches, and snacks
- Full health and dental benefits, including a separate mental health budget
- 100% parental leave top-up for up to 6 months
- Personal enrichment benefits for arts and culture, fitness and well-being, quality time, and workspace improvement
- Remote-flexible work with offices in Toronto, Montreal, San Francisco, New York, Paris, Seoul, and London
- Co-working stipend
- Six weeks of vacation, or 30 working days
About Cohere
Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation models and end-to-end AI products designed to solve real-world business problems. We partner closely with companies to deliver seamless integration, full customization, and easy-to-use solutions for their workforce and customers. Our all-in-one platform offers enterprises the highest levels of data security, privacy and optionality to deploy across all major cloud providers, private cloud environments, or on-premises. HQ: 171 John Street, 2nd Floor, Toronto, ON M5T 1X3
