
Staff Engineer, Distributed Storage and HPC & AI Infrastructure
Together AI8 hours ago
Bengaluru, IndiaStaff+
Responsibilities
- Architect and implement the storage technical strategy and roadmap while driving high-performance architectural decisions for a growing GPU fleet.
- Engineer and scale multi-petabyte AI/ML storage systems using Vast, Weka, and Ceph, including automated tiering and lifecycle policies for cost optimization.
- Design intelligent caching and tiered storage architectures for extreme IOPS and cluster-wide throughput across training and inference workloads.
- Tune L2/L3 storage network isolation to provide secure, production-grade multi-tenancy.
- Develop Kubernetes storage operators and controllers for automated provisioning, self-service abstractions, and quota enforcement.
- Optimize end-to-end data paths, model-weight and dataset caching, parallel filesystems, benchmarking, and profiling to support 10+ GB/s per GPU node and thousands of nodes.
- Contribute high-impact code to open-source storage projects and internal tooling.
Requirements
- 8+ years of storage engineering experience managing distributed storage at multi-petabyte scale.
- Demonstrated experience deploying and operating high-performance storage for GPU or HPC clusters.
- Deep production experience with Kubernetes and cloud-native storage.
- Strong Go and Python programming skills for production systems and tooling.
- BS or MS in Computer Science, Engineering, or equivalent practical experience.
- Technical leadership experience designing systems that materially improved performance, reliability, or cost efficiency.
- Deep expertise with one or more of Ceph, WekaFS, Lustre, Vast, GPFS, or similar parallel filesystems at multi-petabyte scale.
- Production object-storage experience with S3, MinIO, Ceph, or R2, including performance optimization and cost management.
- Experience with Kubernetes storage components including CSI drivers, StatefulSets, PersistentVolumes, storage operators, and custom controllers.
- Experience with GPU storage optimization, RDMA or InfiniBand networking, and parallel filesystem optimization.
- Experience with Terraform, Ansible, Helm, GitOps, and ArgoCD.
- Advanced Linux storage knowledge including ext4, XFS, LVM, NVMe optimization, and RAID configurations.
- Experience operating observability systems including Prometheus, Grafana, and Thanos.
- Nice-to-have experience with GPU Direct Storage, NVMe-oF, storage networking, RDMA implementations, AI/ML storage patterns, and storage benchmarking or profiling tools such as fio, iperf3, iostat, and blktrace.
About Together AI
Together AI builds an AI-native cloud platform for developers, offering high-performance inference, fine-tuning/model shaping, and large-scale pre-training on on-demand GPU clusters with APIs and managed services. It emphasizes open-source models that teams can run and adapt, and also provides infrastructure for decentralized and scalable workloads. Founded in 2022 and headquartered in San Francisco, it is privately held and reports notable customers including Cursor, ElevenLabs, Salesforce, and Zoom.