7 months ago
Responsibilities
- Build and scale Kubernetes-based GPU/TPU superclusters across multiple clouds for high-throughput, low-latency AI workloads
- Collaborate with cloud providers to optimize ML infrastructure for cost, reliability, and performance using RDMA, NCCL, and high-speed interconnects
- Diagnose and resolve infrastructure bottlenecks, performance degradation, and system failures
- Design self-service interfaces and workflows for researchers to monitor, debug, and optimize training jobs
- Translate emerging needs involving JAX, PyTorch, and distributed training into scalable infrastructure solutions
- Champion observability, automation, and infrastructure-as-code practices
- Mentor and collaborate with engineers through code reviews, documentation, and knowledge sharing
- Participate in a compensated 24x7 on-call rotation
Requirements
- Deep expertise with GPU/TPU clusters, distributed training frameworks, and HPC environments
- Experience deploying, managing, and troubleshooting Kubernetes clusters at scale for AI workloads
- Proficiency in Python for ML tooling and Go for systems engineering
- Familiarity with Linux internals, RDMA networking, and performance optimization for ML workloads
- Track record of collaborating with AI researchers or ML engineers on infrastructure challenges
- Ability to identify bottlenecks, propose solutions, and drive impact in a fast-paced environment
- Open-source contribution experience is preferred
Benefits
- Open and inclusive culture and work environment
- Weekly lunch stipend, in-office lunches, and snacks
- Full health and dental benefits with a separate mental health budget
- 100% parental leave top-up for up to 6 months
- Personal enrichment benefits for arts and culture, fitness and well-being, quality time, and workspace improvement
- Remote-flexible work with offices in Toronto, New York, San Francisco, London, and Paris, plus a co-working stipend
- Six weeks of vacation, or 30 working days
- Compensated participation in a 24x7 on-call rotation
Tech Stack
Categories
About Cohere
Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation models and end-to-end AI products designed to solve real-world business problems. We partner closely with companies to deliver seamless integration, full customization, and easy-to-use solutions for their workforce and customers. Our all-in-one platform offers enterprises the highest levels of data security, privacy and optionality to deploy across all major cloud providers, private cloud environments, or on-premises. HQ: 171 John Street, 2nd Floor, Toronto, ON M5T 1X3
