Prime Intellect

Member of Technical Staff - Storage Infrastructure

Prime Intellect
Apply
7 hours ago
Remote, United States or San Francisco, CA, USAMid Level

Base Salary

$150k - $300k/yr

Responsibilities

  • Design and operate storage architectures for training datasets, checkpoints, inference artifacts, and shared research workflows.
  • Deploy and tune parallel filesystems, object storage, and local NVMe caching for demanding AI workloads.
  • Benchmark throughput, latency, metadata performance, and concurrent access using representative workloads.
  • Build provisioning, capacity planning, lifecycle management, and operational automation for storage services.
  • Design and test replication, recovery, backup, and failure-handling procedures against durability and availability targets.
  • Diagnose performance and reliability issues across applications, clients, networks, filesystems, and storage devices.
  • Implement access controls, tenant separation, quotas, monitoring, and runbooks while collaborating with compute and networking teams.

Requirements

  • At least 3 years of experience building or operating production distributed storage systems.
  • Hands-on experience with at least one parallel or distributed filesystem or object storage platform, such as Lustre, BeeGFS, Ceph, or GPFS.
  • Strong Linux administration and performance troubleshooting skills.
  • Experience automating infrastructure operations with Python, Go, Bash, or similar languages.
  • Understanding of storage failure modes, data integrity, consistency, replication, and recovery.
  • Knowledge of block, file, and object storage semantics; NVMe/SSD performance; filesystem tuning; I/O profiling; benchmarking; storage networking; capacity forecasting; observability; alerting; authentication; authorization; encryption; and secure data lifecycle management.
  • Preferred experience includes large GPU training clusters, high-volume checkpoint workloads, S3-compatible object storage, data tiering, distributed caching, RDMA-enabled storage, GPUDirect Storage, Kubernetes storage integrations, SLURM environments, storage cost optimization, or open-source storage contributions.

Benefits

  • Cash compensation of $150,000–$300,000 plus equity incentives.
  • Work directly with customers and engineering teams building large-scale AI infrastructure.

Categories

Prime Intellect

About Prime Intellect

51-200 employees

Prime Intellect builds a full-stack platform for training frontier AI models, giving enterprises access to distributed compute and agentic training infrastructure to develop their own systems. The company operates as a platform and open research lab, with customers leveraging its infrastructure rather than building it in-house. Headquartered in San Francisco and privately held, it has raised a Series A round backed by investors including Founders Fund and NVIDIA.

Contact me