Together AI

Staff Software Engineer, Inference / Compute Infrastructure Engineering

Together AI
Apply
4 hours ago
Amsterdam, NetherlandsStaff+
H1B Sponsor

Responsibilities

  • Design and implement versioned software state machines for physical host and inference-cluster provisioning, lifecycle management, and decommissioning.
  • Build declarative self-service APIs and a control plane that lets engineering teams provision, scale, and tear down inference clusters.
  • Automate self-healing by detecting failed nodes, draining them, triggering repair or replacement, and restoring healthy capacity.
  • Ensure provisioning workflows provide idempotency, retries, rollback, and drift detection in production.
  • Partner with the inference and ML platform teams to model cluster topology, interconnect, scheduling constraints, and software-stack requirements.
  • Deliver typed, tested, versioned infrastructure software through code review and CI/CD, and operate and support it in production.

Requirements

  • Strong software engineering experience with Go, Python, Rust, or a similar language, including writing and testing production software.
  • Experience with durable workflow orchestration tools such as Temporal, Cadence, or equivalent.
  • Experience building software control planes or orchestration systems that model state and continuously reconcile it over time.
  • Experience designing event-driven systems using message queues, event streams, or pub/sub rather than polling or cron-driven scripts.
  • Experience building internal platforms or APIs for other engineering teams with attention to developer experience.
  • Exposure to bare-metal provisioning, PXE/iPXE, Redfish/IPMI, BMC systems, networking fundamentals, VLANs, BGP, fabric design, or GPU and accelerator infrastructure is preferred.
  • Experience with GPU cluster software stacks such as NCCL, CUDA, and InfiniBand/RoCE is preferred.
  • Experience at a hyperscaler, GPU cloud, or datacenter-scale infrastructure organization is preferred.
  • Systems programming experience in Rust or Go is preferred.
Together AI

About Together AI

201-500 employees

Together AI is the AI Native Cloud, purpose-built for AI engineers and researchers with a full suite of tooling across inference, model shaping, and pre-training. AI natives can use Together AI as a full-stack AI platform — from a high- performance inference engine built for reliable and fast scaling to on-demand GPU clusters and massive-scale AI factories. Together AI continuously pushes the frontier forward by productizing cutting-edge research from our world-leading AI systems research team. By combining research velocity with production-grade infrastructure, we enable companies to reliably scale AI-native applications as fast as the field evolves. Trusted by leading AI natives like Cursor, Decagon, Eleven Labs, AI21, Hedra, and Cartesia, as well as SaaS innovators such as Salesforce, Zoom, and Zomato, Together AI powers the next generation of AI-native applications.