
Member of Technical Staff - Storage Infrastructure
Prime Intellect7 hours ago
Remote, United States or San Francisco, CA, USAMid Level
Base Salary
$150k - $300k/yr
Responsibilities
- Design and operate storage architectures for training datasets, checkpoints, inference artifacts, and shared research workflows.
- Deploy and tune parallel filesystems, object storage, and local NVMe caching for demanding AI workloads.
- Benchmark throughput, latency, metadata performance, and concurrent access using representative workloads.
- Build provisioning, capacity planning, lifecycle management, and operational automation for storage services.
- Design and test replication, recovery, backup, and failure-handling procedures against durability and availability targets.
- Diagnose performance and reliability issues across applications, clients, networks, filesystems, and storage devices.
- Implement access controls, tenant separation, quotas, monitoring, and runbooks while collaborating with compute and networking teams.
Requirements
- At least 3 years of experience building or operating production distributed storage systems.
- Hands-on experience with at least one parallel or distributed filesystem or object storage platform, such as Lustre, BeeGFS, Ceph, or GPFS.
- Strong Linux administration and performance troubleshooting skills.
- Experience automating infrastructure operations with Python, Go, Bash, or similar languages.
- Understanding of storage failure modes, data integrity, consistency, replication, and recovery.
- Knowledge of block, file, and object storage semantics; NVMe/SSD performance; filesystem tuning; I/O profiling; benchmarking; storage networking; capacity forecasting; observability; alerting; authentication; authorization; encryption; and secure data lifecycle management.
- Preferred experience includes large GPU training clusters, high-volume checkpoint workloads, S3-compatible object storage, data tiering, distributed caching, RDMA-enabled storage, GPUDirect Storage, Kubernetes storage integrations, SLURM environments, storage cost optimization, or open-source storage contributions.
Benefits
- Cash compensation of $150,000–$300,000 plus equity incentives.
- Work directly with customers and engineering teams building large-scale AI infrastructure.
Tech Stack
Categories
About Prime Intellect
Prime Intellect builds a full-stack platform for training frontier AI models, giving enterprises access to distributed compute and agentic training infrastructure to develop their own systems. The company operates as a platform and open research lab, with customers leveraging its infrastructure rather than building it in-house. Headquartered in San Francisco and privately held, it has raised a Series A round backed by investors including Founders Fund and NVIDIA.