7 months ago
Responsibilities
- Build self-service systems that automate service management, deployment, and operations
- Develop and maintain custom Kubernetes operators for language model deployments
- Automate environment observability and resilience and enable developers to troubleshoot and resolve problems
- Help ensure defined SLOs are met, including participating in an on-call rotation
- Build relationships with internal developers and influence the Infrastructure team roadmap
- Contribute to team development through knowledge sharing and active reviews
- Deploy optimized NLP models to low-latency, high-throughput, highly available production environments
- Interface with customers and create customized deployments for their needs
Requirements
- 5+ years of engineering experience running production infrastructure at large scale
- Experience designing highly available distributed systems with Kubernetes and GPU workloads
- Experience with Kubernetes development, production coding, and support
- Experience with GCP, Azure, AWS, OCI, and multi-cloud, on-premises, or hybrid serving environments
- Experience designing, deploying, supporting, and troubleshooting complex Linux-based computing environments
- Experience managing compute, storage, network resources, and costs
- Familiarity with GPUs, TPUs, and custom accelerators and their effects on inference latency and throughput
- Strong understanding of or working experience with distributed systems
- Experience with Golang, C++, or other languages designed for high-performance scalable servers
- Excellent collaboration and troubleshooting skills, plus adaptability in solving evolving technical challenges
Benefits
- Open and inclusive culture and work environment
- Weekly lunch stipend, in-office lunches, and snacks
- Full health and dental benefits, including a separate mental health budget
- 100% parental leave top-up for up to 6 months
- Personal enrichment benefits for arts and culture, fitness and well-being, quality time, and workspace improvement
- Remote-flexible work, offices in Toronto, New York, San Francisco, London, and Paris, plus a co-working stipend
- Six weeks of vacation, or 30 working days
Categories
DevOpsSite Reliability
About Cohere
Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation models and end-to-end AI products designed to solve real-world business problems. We partner closely with companies to deliver seamless integration, full customization, and easy-to-use solutions for their workforce and customers. Our all-in-one platform offers enterprises the highest levels of data security, privacy and optionality to deploy across all major cloud providers, private cloud environments, or on-premises. HQ: 171 John Street, 2nd Floor, Toronto, ON M5T 1X3
