Cohere

Member of Technical Staff, Model Efficiency

Cohere
Apply
11 months ago
Toronto, Canada +3 moreSenior
H1B sponsor

Responsibilities

  • Improve LLM inference performance across the model execution stack.
  • Identify bottlenecks and develop optimizations that reduce latency and increase throughput.
  • Collaborate with modeling and systems teams to experiment, measure, and ship inference improvements.
  • Build expertise in GPU/CUDA optimization, kernel-level improvements, and model execution strategies for MoE and large-scale architectures.

Requirements

  • At least 5 years of experience writing high-performance, production-quality code.
  • Strong programming skills in C++ or Python; Rust and Go are also welcome.
  • Experience working with large language models and familiarity with the LLM inference ecosystem, including vLLM or SGLang.
  • Ability to diagnose and resolve performance bottlenecks across the model execution stack.
  • Experience with GPU programming, CUDA, or low-level systems optimization is a plus.
  • Experience with transformer language modeling, including MoE, speculative decoding, or KV-cache optimizations, is a plus.
  • Experience scaling performance-critical distributed systems for computation, search, or storage is a plus.

Benefits

  • Open and inclusive culture and work environment
  • Weekly lunch stipend, in-office lunches, and snacks
  • Full health and dental benefits, including a separate mental health budget
  • 100% parental leave top-up for up to 6 months
  • Personal enrichment benefits for arts and culture, fitness and well-being, quality time, and workspace improvement
  • Remote-flexible work with offices in Toronto, Montreal, San Francisco, New York, Paris, Seoul, and London
  • Co-working stipend
  • Six weeks of vacation, or 30 working days
Cohere

About Cohere

1,001-5,000 employees

Cohere builds large language models and an enterprise AI platform that companies use for search, summarization, and workflow automation, delivered via API or private deployments. Founded in 2019 and headquartered in Toronto, it focuses on multilingual models, data controls, and options to run across major clouds or on-premises. The business is privately held and serves security- and compliance-sensitive organizations.

Contact me