Together AI

Senior Software Engineer - Together Cloud Infrastructure

Together AI
Apply
4 months ago

Base Salary

$160k - $230k/yr

Responsibilities

  • Design, build, and maintain secure, performant, highly available backend services and operators for data-center hardware management.
  • Build the IaaS software layer for a new GB200 data center containing thousands of GPUs.
  • Develop and operate a global, multi-exabyte high-performance object store for large pretraining datasets.
  • Build customer observability stacks and automate node lifecycle management for fault-tolerant distributed pretraining.
  • Perform architecture and research work for decentralized AI workloads.
  • Contribute to the open-source Together AI platform and create services, tools, and developer documentation.
  • Create testing frameworks focused on robustness and fault tolerance.

Requirements

  • At least five years of professional software development experience and proficiency in at least one backend programming language, with Golang desired.
  • At least five years of experience writing high-performance, well-tested, production-quality code.
  • Experience building and operating high-performance or globally distributed microservice architectures across AWS, Azure, or GCP.
  • Strong systems knowledge and troubleshooting abilities across compute, networking, and storage, including concurrency, memory management, performant I/O, and scale.
  • Experience with infrastructure automation tools such as Terraform and Ansible, monitoring and observability stacks such as Prometheus and Grafana, and CI/CD pipelines such as GitHub Actions and ArgoCD.
  • Kubernetes internals experience, including operators, device/storage/network plugins, custom schedulers, or Kubernetes patches, is a strong plus.
  • Experience with VMs and hypervisors such as QEMU/KVM, cloud-hypervisor, VFIO, virtio, PCIe passthrough, KubeVirt, and SR-IOV is a strong plus.
  • Experience with data-center networking technologies and solutions such as VLAN, VXLAN, VPN, VPC, OVS, and OVN is a strong plus.
  • Experience with Cluster API or similar systems, high-performance computing, networking or storage, GPU or InfiniBand virtualization, IaaS or PaaS systems, DPUs or SmartNICs, NCCL, or CUDA is advantageous.
  • Excellent communication skills, including the ability to write clear design documents and work with technical and non-technical stakeholders.

Benefits

  • Competitive compensation, startup equity, health insurance, and other benefits are offered.
  • Flexible remote-work arrangements are available.
  • This is a full-time position.

Tech Stack

AnsibleAWSGitHub ActionsGoGoogle Cloud PlatformGrafanaKubernetesPrometheusTerraform

Categories

Together AI

About Together AI

201-500 employees

Together AI is the AI Native Cloud, purpose-built for AI engineers and researchers with a full suite of tooling across inference, model shaping, and pre-training. AI natives can use Together AI as a full-stack AI platform — from a high- performance inference engine built for reliable and fast scaling to on-demand GPU clusters and massive-scale AI factories. Together AI continuously pushes the frontier forward by productizing cutting-edge research from our world-leading AI systems research team. By combining research velocity with production-grade infrastructure, we enable companies to reliably scale AI-native applications as fast as the field evolves. Trusted by leading AI natives like Cursor, Decagon, Eleven Labs, AI21, Hedra, and Cartesia, as well as SaaS innovators such as Salesforce, Zoom, and Zomato, Together AI powers the next generation of AI-native applications.

Contact me