
AI infrastructure Engineer (SRE) Amsterdam
Together AI5 months ago
Amsterdam, NetherlandsSenior
Responsibilities
- Participate in a PagerDuty on-call rotation and respond to incidents affecting availability.
- Build and operate infrastructure with Ansible, Terraform, and Kubernetes to support massive concurrent-user scale.
- Build monitoring systems to maintain high-quality customer service.
- Design and implement operational processes such as deployments and upgrades.
- Debug production issues across all services and levels of the stack.
- Identify product architecture improvements related to reliability, performance, and availability.
- Plan the growth of Together AI’s infrastructure.
Requirements
- 7+ years of professional SRE or related experience.
- Bachelor's degree in Computer Science or a related field, or equivalent work experience.
- Expert knowledge of Ansible, including roles and playbooks, Terraform, and Kubernetes.
- Proficiency in programming and scripting languages.
- Direct experience with monitoring and observability practices.
- Advanced knowledge of cloud services.
- Ability to collaborate with different stakeholders and subject matter experts.
Tech Stack
Categories
DevOpsSite Reliability
About Together AI
Together AI builds an AI-native cloud platform for developers, offering high-performance inference, fine-tuning/model shaping, and large-scale pre-training on on-demand GPU clusters with APIs and managed services. It emphasizes open-source models that teams can run and adapt, and also provides infrastructure for decentralized and scalable workloads. Founded in 2022 and headquartered in San Francisco, it is privately held and reports notable customers including Cursor, ElevenLabs, Salesforce, and Zoom.