Together AI

Staff Software Engineer, GPU Infrastructure Lifecycle Management

Together AI
Apply
2 months ago

Base Salary

$240k - $280k/yr

Responsibilities

  • Design and implement versioned software state machines for the full physical-host lifecycle, from discovery and GPU driver/CUDA bring-up through health validation and decommissioning or RMA.
  • Build declarative self-service APIs and a control plane that let inference teams provision, scale, and tear down clusters without manual intervention.
  • Automate detection, draining, repair or replacement, and reintegration of degraded or failed nodes.
  • Ensure pipeline reliability through idempotency, retries, rollback, and drift detection.
  • Partner with inference and ML platform teams to encode cluster shapes, topology, interconnects, and scheduling constraints as platform abstractions.
  • Operate and support the production platform while maintaining strong typing, automated tests, versioning, code review, and deployment practices.

Requirements

  • Strong software engineering background in Go, Python, Rust, or similar, including writing and testing production software.
  • Experience with durable workflow orchestration tools such as Temporal, Cadence, or equivalent.
  • Experience building software control planes or orchestration systems that model state and continuously reconcile it, such as Kubernetes controllers or operators.
  • Experience designing event-driven systems using message queues, event streams, or pub/sub rather than polling or cron-driven scripts.
  • Experience building internal platforms or APIs consumed by other engineering teams with attention to developer experience.
  • Exposure to bare-metal provisioning technologies, networking fundamentals, or GPU and accelerator infrastructure is preferred.
  • Experience with GPU cluster software stacks such as NCCL, CUDA, and InfiniBand/RoCE is preferred.
  • Prior experience at a hyperscaler, GPU cloud, or datacenter-scale infrastructure organization is preferred.
  • Systems programming experience in Rust or Go is preferred.

Benefits

  • Startup equity
  • Health insurance
  • Other competitive benefits
  • Full-time position in the US
Together AI

About Together AI

201-500 employees

Together AI is the AI Native Cloud, purpose-built for AI engineers and researchers with a full suite of tooling across inference, model shaping, and pre-training. AI natives can use Together AI as a full-stack AI platform — from a high- performance inference engine built for reliable and fast scaling to on-demand GPU clusters and massive-scale AI factories. Together AI continuously pushes the frontier forward by productizing cutting-edge research from our world-leading AI systems research team. By combining research velocity with production-grade infrastructure, we enable companies to reliably scale AI-native applications as fast as the field evolves. Trusted by leading AI natives like Cursor, Decagon, Eleven Labs, AI21, Hedra, and Cartesia, as well as SaaS innovators such as Salesforce, Zoom, and Zomato, Together AI powers the next generation of AI-native applications.

Contact me