Together AI

Staff Engineer, Distributed Storage and HPC & AI Infrastructure

Together AI
Apply
3 months ago

Base Salary

$250k - $300k/yr

Responsibilities

  • Design multi-petabyte AI/ML storage systems using technologies such as WekaFS, Ceph, and Lustre, while leading capacity planning and cost optimization.
  • Design and optimize RDMA, InfiniBand, 400GbE, NVMe-oF, iSCSI, and TCP/IP storage data paths for high throughput and low latency.
  • Build Kubernetes storage operators and controllers for automated provisioning, self-service abstractions, multi-tenant isolation, quotas, and reusable Helm/Terraform patterns.
  • Optimize caching, data locality, prefetching, eviction, model-weight distribution, parallel filesystems, and data paths at thousands-of-node scale.
  • Implement monitoring, alerting, SLOs, disaster recovery, backups, runbooks, chaos engineering, and automated remediation for 99.9%+ uptime.
  • Partner with ML and SRE teams, mentor others on storage practices, contribute to open source, and write documentation, postmortems, and public technical learnings.

Requirements

  • 8+ years of storage engineering experience, including 3+ years managing distributed storage at multi-petabyte scale.
  • Proven experience deploying and operating high-performance storage for GPU/HPC clusters in production.
  • Deep Kubernetes and cloud-native storage experience.
  • Strong production-grade coding skills in Go and Python.
  • A BS or MS in Computer Science, Engineering, or equivalent practical experience.
  • A track record of technical leadership and delivering significant performance, reliability, or cost improvements.
  • Deep expertise with parallel filesystems such as WekaFS, Lustre, GPFS, or BeeGFS at multi-petabyte scale.
  • Production experience with object storage such as S3, MinIO, Ceph, or R2, including performance optimization and cost management.
  • Experience with Kubernetes storage components including CSI drivers, StatefulSets, PersistentVolumes, storage operators, and custom controllers.
  • Experience optimizing storage for GPU workloads and RDMA/InfiniBand networking, including parallel filesystem optimization.
  • Experience with Terraform, Ansible, Helm, and GitOps using ArgoCD.
  • Advanced knowledge of Linux storage, ext4, XFS, LVM, NVMe optimization, and RAID configurations.
  • Experience operating observability systems with Prometheus, Grafana, and Thanos.
  • Nice-to-have experience includes GPU Direct Storage, NVMe-oF, 100GbE/400GbE storage networking, Kubernetes operator development, storage snapshots and cloning, thin provisioning, backup and disaster recovery, storage encryption, and storage benchmarking tools.

Benefits

  • Remote-work flexibility is offered.
  • Health insurance, startup equity, and other benefits are provided.
  • This is a full-time position.

Tech Stack

AnsibleGoGrafanaHelmKubernetesLinuxPrometheusPythonTerraform

Categories

Together AI

About Together AI

201-500 employees

Together AI is the AI Native Cloud, purpose-built for AI engineers and researchers with a full suite of tooling across inference, model shaping, and pre-training. AI natives can use Together AI as a full-stack AI platform — from a high- performance inference engine built for reliable and fast scaling to on-demand GPU clusters and massive-scale AI factories. Together AI continuously pushes the frontier forward by productizing cutting-edge research from our world-leading AI systems research team. By combining research velocity with production-grade infrastructure, we enable companies to reliably scale AI-native applications as fast as the field evolves. Trusted by leading AI natives like Cursor, Decagon, Eleven Labs, AI21, Hedra, and Cartesia, as well as SaaS innovators such as Salesforce, Zoom, and Zomato, Together AI powers the next generation of AI-native applications.

Contact me