Together AI

AI infrastructure Engineer (SRE) Amsterdam

Together AI
Apply
5 months ago
Amsterdam, NetherlandsSenior

Responsibilities

  • Participate in a PagerDuty on-call rotation and respond to incidents affecting availability.
  • Build and operate infrastructure with Ansible, Terraform, and Kubernetes to support massive concurrent-user scale.
  • Build monitoring systems to maintain high-quality customer service.
  • Design and implement operational processes such as deployments and upgrades.
  • Debug production issues across all services and levels of the stack.
  • Identify product architecture improvements related to reliability, performance, and availability.
  • Plan the growth of Together AI’s infrastructure.

Requirements

  • 7+ years of professional SRE or related experience.
  • Bachelor's degree in Computer Science or a related field, or equivalent work experience.
  • Expert knowledge of Ansible, including roles and playbooks, Terraform, and Kubernetes.
  • Proficiency in programming and scripting languages.
  • Direct experience with monitoring and observability practices.
  • Advanced knowledge of cloud services.
  • Ability to collaborate with different stakeholders and subject matter experts.

Tech Stack

Categories

DevOpsSite Reliability
Together AI

About Together AI

201-500 employees

Together AI builds an AI-native cloud platform for developers, offering high-performance inference, fine-tuning/model shaping, and large-scale pre-training on on-demand GPU clusters with APIs and managed services. It emphasizes open-source models that teams can run and adapt, and also provides infrastructure for decentralized and scalable workloads. Founded in 2022 and headquartered in San Francisco, it is privately held and reports notable customers including Cursor, ElevenLabs, Salesforce, and Zoom.

Contact me