
Cloud Storage Integration Engineer (Storage / Image / Registry)
BitDeer Technologies Group11 days ago
Remote, United States or San Jose, CA, USAMid Level / Senior
Responsibilities
- Design and integrate distributed and parallel file systems into the GPU cloud, optimizing for AI training and inference I/O patterns.
- Own storage provisioning, mounting, multi-tenant isolation, quota management, and lifecycle integration within the platform control plane.
- Tune throughput and latency for parallel access, including dataset loading and checkpointing, and benchmark across GPU SKUs and workloads.
- Architect multi-region storage covering data locality, replication, consistency, durability, failure domains, and cross-region access.
- Own golden images, templates, GPU drivers, CUDA, and the container/image registry, including versioned releases and multi-region distribution.
- Build monitoring, capacity-planning processes, and runbooks, and eliminate single points of failure.
- Partner with Compute, Network, and Control Plane teams on delivery, storage networking, provisioning, and quota capabilities.
Requirements
- 3+ years of experience in storage engineering or platform infrastructure; senior-level candidates require 6+ years.
- Hands-on experience with distributed or parallel file systems and production operation of systems such as Ceph, Lustre, GPFS/Spectrum Scale, BeeGFS, JuiceFS, or MinIO.
- Strong understanding of distributed file-system internals, including data and metadata separation, replication, consistency models, and POSIX versus object semantics.
- Experience tuning high-throughput parallel I/O and familiarity with NVMe, RDMA/RoCE storage networking, and caching.
- Strong Linux systems expertise and automation skills using Python or Go.
- HPC, AI storage, or multi-region storage experience is preferred.