
Staff Software Engineer, Inference / Compute Infrastructure Engineering
Together AIResponsibilities
- Design and implement versioned provisioning state machines for the full lifecycle of physical hosts and GPU clusters.
- Build declarative self-service APIs and a control plane for requesting, scaling, and tearing down inference clusters.
- Automate self-healing by detecting degraded nodes, draining them, triggering repair or replacement, and restoring healthy capacity.
- Own pipeline reliability through idempotency, retries, rollback, and drift detection.
- Partner with inference and ML platform teams to encode cluster shapes, topology, interconnects, and scheduling constraints.
- Develop typed, tested, versioned infrastructure software and operate and support it in production.
Requirements
- Strong software engineering background in Go, Python, Rust, or a similar language, including writing and testing production software.
- Experience with durable workflow orchestration tools such as Temporal, Cadence, or equivalent.
- Experience building control planes or orchestration systems that model state and reconcile it over time, such as Kubernetes controllers or operators.
- Experience designing event-driven systems using message queues, event streams, or pub/sub systems such as Kafka, NATS, or SQS.
- Experience building internal platforms or APIs for other engineering teams with attention to developer experience.
- Preferred exposure to bare-metal provisioning technologies such as PXE, iPXE, Redfish, IPMI, or BMC; networking fundamentals; or GPU and accelerator infrastructure.
- Preferred experience with GPU cluster software stacks such as NCCL, CUDA, and InfiniBand or RoCE.
- Prior experience at a hyperscaler, GPU cloud, or datacenter-scale infrastructure organization is preferred.
- Systems programming experience in Rust or Go is preferred.
Benefits
- The role is part of Together AI, a research-driven company focused on open and transparent AI systems and AI infrastructure.
- The company offers equal employment opportunity regardless of protected characteristics.
Tech Stack
About Together AI
Together AI is the AI Native Cloud, purpose-built for AI engineers and researchers with a full suite of tooling across inference, model shaping, and pre-training. AI natives can use Together AI as a full-stack AI platform — from a high- performance inference engine built for reliable and fast scaling to on-demand GPU clusters and massive-scale AI factories. Together AI continuously pushes the frontier forward by productizing cutting-edge research from our world-leading AI systems research team. By combining research velocity with production-grade infrastructure, we enable companies to reliably scale AI-native applications as fast as the field evolves. Trusted by leading AI natives like Cursor, Decagon, Eleven Labs, AI21, Hedra, and Cartesia, as well as SaaS innovators such as Salesforce, Zoom, and Zomato, Together AI powers the next generation of AI-native applications.