
Staff Software Engineer, Inference / Compute Infrastructure Engineering
Together AI2 months ago
Amsterdam, NetherlandsStaff+
Responsibilities
- Design and implement versioned software state machines for physical host and inference-cluster provisioning, lifecycle management, and decommissioning.
- Build declarative self-service APIs and a control plane that lets engineering teams provision, scale, and tear down inference clusters.
- Automate self-healing by detecting failed nodes, draining them, triggering repair or replacement, and restoring healthy capacity.
- Ensure provisioning workflows provide idempotency, retries, rollback, and drift detection in production.
- Partner with the inference and ML platform teams to model cluster topology, interconnect, scheduling constraints, and software-stack requirements.
- Deliver typed, tested, versioned infrastructure software through code review and CI/CD, and operate and support it in production.
Requirements
- Strong software engineering experience with Go, Python, Rust, or a similar language, including writing and testing production software.
- Experience with durable workflow orchestration tools such as Temporal, Cadence, or equivalent.
- Experience building software control planes or orchestration systems that model state and continuously reconcile it over time.
- Experience designing event-driven systems using message queues, event streams, or pub/sub rather than polling or cron-driven scripts.
- Experience building internal platforms or APIs for other engineering teams with attention to developer experience.
- Exposure to bare-metal provisioning, PXE/iPXE, Redfish/IPMI, BMC systems, networking fundamentals, VLANs, BGP, fabric design, or GPU and accelerator infrastructure is preferred.
- Experience with GPU cluster software stacks such as NCCL, CUDA, and InfiniBand/RoCE is preferred.
- Experience at a hyperscaler, GPU cloud, or datacenter-scale infrastructure organization is preferred.
- Systems programming experience in Rust or Go is preferred.
Tech Stack
About Together AI
Together AI builds an AI-native cloud platform for developers, offering high-performance inference, fine-tuning/model shaping, and large-scale pre-training on on-demand GPU clusters with APIs and managed services. It emphasizes open-source models that teams can run and adapt, and also provides infrastructure for decentralized and scalable workloads. Founded in 2022 and headquartered in San Francisco, it is privately held and reports notable customers including Cursor, ElevenLabs, Salesforce, and Zoom.