
Staff Software Engineer, Inference / Compute Infrastructure Engineering
Together AI2 months ago
London, United KingdomStaff+
Responsibilities
- Design and implement versioned provisioning state machines for the full lifecycle of physical hosts and GPU clusters.
- Build declarative self-service APIs and a control plane for requesting, scaling, and tearing down inference clusters.
- Automate self-healing by detecting degraded nodes, draining them, triggering repair or replacement, and restoring healthy capacity.
- Own pipeline reliability through idempotency, retries, rollback, and drift detection.
- Partner with inference and ML platform teams to encode cluster shapes, topology, interconnects, and scheduling constraints.
- Develop typed, tested, versioned infrastructure software and operate and support it in production.
Requirements
- Strong software engineering background in Go, Python, Rust, or a similar language, including writing and testing production software.
- Experience with durable workflow orchestration tools such as Temporal, Cadence, or equivalent.
- Experience building control planes or orchestration systems that model state and reconcile it over time, such as Kubernetes controllers or operators.
- Experience designing event-driven systems using message queues, event streams, or pub/sub systems such as Kafka, NATS, or SQS.
- Experience building internal platforms or APIs for other engineering teams with attention to developer experience.
- Preferred exposure to bare-metal provisioning technologies such as PXE, iPXE, Redfish, IPMI, or BMC; networking fundamentals; or GPU and accelerator infrastructure.
- Preferred experience with GPU cluster software stacks such as NCCL, CUDA, and InfiniBand or RoCE.
- Prior experience at a hyperscaler, GPU cloud, or datacenter-scale infrastructure organization is preferred.
- Systems programming experience in Rust or Go is preferred.
Benefits
- The role is part of Together AI, a research-driven company focused on open and transparent AI systems and AI infrastructure.
- The company offers equal employment opportunity regardless of protected characteristics.
Tech Stack
About Together AI
Together AI builds an AI-native cloud platform for developers, offering high-performance inference, fine-tuning/model shaping, and large-scale pre-training on on-demand GPU clusters with APIs and managed services. It emphasizes open-source models that teams can run and adapt, and also provides infrastructure for decentralized and scalable workloads. Founded in 2022 and headquartered in San Francisco, it is privately held and reports notable customers including Cursor, ElevenLabs, Salesforce, and Zoom.