
Senior AI Storage Infrastructure Engineer
BitDeer Technologies Group14 days ago
Remote, United States or San Jose, CA, USASenior
Responsibilities
- Design, deploy, and maintain Kubernetes Container Storage Interface drivers for high-performance parallel file systems such as Weka, Lustre, DAOS, and VAST.
- Architect and implement GPUDirect Storage integrations for direct data transfers between NVMe drives and GPU memory.
- Develop local NVMe caching strategies for model weights and datasets used in distributed training.
- Optimize IOPS, throughput, and latency across the containerized storage stack.
- Work with the GPU Systems & Fabric team to optimize storage for RDMA, InfiniBand, and RoCE interconnects.
- Implement monitoring and alerting to detect storage contention and hardware degradation.
- Define Kubernetes storage policies, quotas, and multi-tenancy isolation strategies.
- Mentor junior engineers and lead infrastructure architecture reviews.
Requirements
- Bachelor’s or master’s degree in Computer Science, Electrical Engineering, or a related field.
- At least 5 years of experience with distributed storage systems and high-performance file systems.
- Deep understanding of POSIX compliance and file I/O semantics.
- Expertise with the Kubernetes CSI paradigm, including volume plugins and storage operators.
- Hands-on experience with block and file I/O at the Linux OS level and kernel-level performance tuning.
- Familiarity with RDMA, InfiniBand, RoCE, and their interaction with storage subsystems.
- Experience operating, debugging, and scaling large-scale storage environments in production or HPC settings.
- Experience with Terraform, Ansible, and CI/CD pipelines.
- Strong technical communication skills and the ability to influence cross-functional architecture decisions.
- Experience in high-velocity, high-growth engineering environments is strongly preferred.