9 months ago
Responsibilities
- Build self-service systems that automate service management, deployment, and operations
- Develop and maintain custom Kubernetes operators for language model deployments
- Automate environment observability and resilience and enable developers to troubleshoot and resolve problems
- Help ensure defined SLOs are met, including participating in an on-call rotation
- Build relationships with internal developers and influence the Infrastructure team roadmap
- Contribute to team development through knowledge sharing and active reviews
- Deploy optimized NLP models to low-latency, high-throughput, highly available production environments
- Interface with customers and create customized deployments for their needs
Requirements
- 5+ years of engineering experience running production infrastructure at large scale
- Experience designing highly available distributed systems with Kubernetes and GPU workloads
- Experience with Kubernetes development, production coding, and support
- Experience with GCP, Azure, AWS, OCI, and multi-cloud, on-premises, or hybrid serving environments
- Experience designing, deploying, supporting, and troubleshooting complex Linux-based computing environments
- Experience managing compute, storage, network resources, and costs
- Familiarity with GPUs, TPUs, and custom accelerators and their effects on inference latency and throughput
- Strong understanding of or working experience with distributed systems
- Experience with Golang, C++, or other languages designed for high-performance scalable servers
- Excellent collaboration and troubleshooting skills, plus adaptability in solving evolving technical challenges
Benefits
- Open and inclusive culture and work environment
- Weekly lunch stipend, in-office lunches, and snacks
- Full health and dental benefits, including a separate mental health budget
- 100% parental leave top-up for up to 6 months
- Personal enrichment benefits for arts and culture, fitness and well-being, quality time, and workspace improvement
- Remote-flexible work, offices in Toronto, New York, San Francisco, London, and Paris, plus a co-working stipend
- Six weeks of vacation, or 30 working days
Categories
DevOpsSite Reliability
About Cohere
Cohere builds large language models and an enterprise AI platform that companies use for search, summarization, and workflow automation, delivered via API or private deployments. Founded in 2019 and headquartered in Toronto, it focuses on multilingual models, data controls, and options to run across major clouds or on-premises. The business is privately held and serves security- and compliance-sensitive organizations.
