1 month ago
Remote, United StatesStaff+
Responsibilities
- Lead the design and implementation of high-performance data movement pipelines using NVIDIA NIXL across GPU, CPU, and storage tiers.
- Architect and integrate DDN Infinia with GPU-accelerated inference platforms for large-scale real-time AI workloads.
- Optimize I/O paths between GPU memory and storage using NVIDIA GPUDirect Storage, RDMA, and NVMe-over-Fabrics.
- Define multi-tier storage architectures using NVMe, SSD, and object storage.
- Develop KV cache offloading, prefetching, and persistence strategies across distributed storage layers.
- Partner with AI/ML teams to optimize inference performance in PyTorch and TensorFlow.
- Establish benchmarking frameworks and lead storage and data-movement performance tuning.
- Diagnose bottlenecks across storage, networking, and GPU subsystems.
- Influence distributed inference architecture for scalability, resilience, and data locality.
- Drive observability, performance monitoring, automation, reliability engineering, and engineering best practices.
- Mentor junior engineers and provide technical leadership across cross-functional teams.
Requirements
- Bachelor’s or Master’s degree in Computer Science, Engineering, or a related field.
- 12+ years of experience in storage systems, distributed systems, or performance engineering.
- Proven experience architecting and delivering large-scale, high-performance infrastructure systems.
- Deep expertise in distributed storage architectures, including object storage, scalable file systems, or cloud-native storage platforms.
- Strong understanding of the Linux I/O stack, filesystem internals, and storage protocols.
- Extensive hands-on experience with NVMe, SSD optimization, and high-performance storage environments.
- Strong experience with RDMA, InfiniBand, or other high-speed data transfer technologies.
- Solid understanding of GPU computing and CPU–GPU data movement patterns.
- Proficiency in Python and/or C/C++, with advanced debugging, profiling, and performance tuning skills.
- Demonstrated ability to optimize latency-sensitive, high-throughput production systems.
- Preferred experience with NVIDIA NIXL or similar data-movement frameworks and GPU-aware storage pipelines.
- Preferred understanding of AI inference systems, LLM serving architectures, KV cache optimization, Retrieval-Augmented Generation pipelines, and open vector search ecosystems.
- Preferred background in high-performance computing or hyperscale distributed environments, caching, memory tiering, data locality, and disaggregated compute and storage architectures.