1 month ago
Santa Clara, CA, USAStaff+
Responsibilities
- Design, develop, and optimize high-performance storage I/O paths for throughput, latency, and concurrency.
- Build distributed storage components for erasure coding, data protection, recovery, replication, and data layout.
- Develop scalable concurrency models, locking mechanisms, and fault-tolerant designs.
- Analyze, profile, benchmark, and tune performance across CPU, memory, caching, scheduling, data movement, multi-node, and multi-device environments.
- Design asynchronous, event-driven architectures supporting AI workloads, high-speed data pipelines, and real-time analytics.
- Lead architecture decisions, design reviews, code reviews, complex debugging, and technical strategy for the I/O Path team.
- Mentor engineers and promote engineering excellence through technical reviews and strong engineering practices.
- Collaborate with QE, storage, networking, performance, field, and support teams to deliver reliable enterprise software.
- Improve validation strategies, quality coverage, automation, release processes, and software delivery reliability.
- Debug customer and system issues and translate findings into product improvements.
Requirements
- 15+ years of experience developing large-scale systems software using C/C++.
- 10+ years of hands-on experience designing and developing storage systems, distributed storage platforms, or high-performance I/O subsystems.
- Deep knowledge of storage architecture, I/O optimization, memory management, caching, scheduling, and data path design.
- Experience with erasure coding, replication, recovery, fault tolerance, concurrency, synchronization, and distributed systems.
- Hands-on experience with SPDK or similar high-performance user-space storage frameworks.
- Expertise in performance analysis, profiling, debugging, and system optimization.
- Preferred experience with HPC, AI infrastructure, large-scale storage platforms, distributed file systems, object storage, or hybrid storage architectures.
- Preferred knowledge of NVMe-oF, RDMA, and high-speed networking technologies.
- Preferred experience building observability, monitoring, and self-healing infrastructure.
- Preferred familiarity with containerized workloads and cluster orchestration.