
Staff Software Engineer - AI Compute, Together Cloud
Together AI5 hours ago
Base Salary
$260k - $300k/yr
Responsibilities
- Own the GPU and network virtualization stack, including hypervisor, kernel, and software-defined networking work.
- Architect and roadmap the in-data-center IaaS layer for compute, storage, networks, VMs, parallel filesystems, VPCs, and InfiniBand partitions.
- Lead the build-out of a Vera Rubin data center from hardware bring-up through customer-facing APIs.
- Design GPU scheduling and global management control planes for on-demand and reserved clusters across multiple data centers.
- Architect monitoring, automated detection, isolation, and recovery systems for infrastructure fault tolerance.
- Set technical direction through design reviews, architectural decisions, engineering standards, and cross-team alignment.
- Mentor engineers, grow team expertise, support hiring, and improve testing frameworks, tools, and documentation.
Requirements
- 7+ years of professional software development experience.
- Expert proficiency in at least one backend programming language, with Golang preferred, and experience writing high-performance, well-tested production code.
- Experience owning the architecture of large distributed systems from initial design through production at scale.
- Deep experience building and operating globally distributed, high-performance microservice architectures across AWS, Azure, or GCP.
- Expert knowledge of compute, networking, and storage systems, including concurrency, memory management, performant I/O, and global scale.
- Demonstrated technical leadership through mentoring, design reviews, and alignment across teams without direct reporting relationships.
- Experience operating reliable customer-facing production systems at scale, including infrastructure automation with Terraform and Ansible, observability with Prometheus and Grafana, and CI/CD with GitHub Actions and ArgoCD.
- Preferred qualifications include Kubernetes internals and operators, VMs and hypervisors, data-center networking, Cluster API, high-performance computing, GPU or InfiniBand virtualization, IaaS or PaaS, DPUs or SmartNICs, and GPU programming with NCCL or CUDA.
Benefits
- US base salary range of $260,000-$300,000 plus equity and benefits.
- Health insurance and other benefits.
- Flexible remote-work arrangements.
- Full-time position.
Tech Stack
About Together AI
Together AI builds an AI-native cloud platform for developers, offering high-performance inference, fine-tuning/model shaping, and large-scale pre-training on on-demand GPU clusters with APIs and managed services. It emphasizes open-source models that teams can run and adapt, and also provides infrastructure for decentralized and scalable workloads. Founded in 2022 and headquartered in San Francisco, it is privately held and reports notable customers including Cursor, ElevenLabs, Salesforce, and Zoom.