3 months ago
Base Salary
$314k - $465k/yr
Responsibilities
- Set technical direction for storage software architecture across the Infrastructure Engineering organization and influence petabyte-scale deployments.
- Author and review design documents for storage systems, protocols, and integrations while raising the team’s technical bar.
- Mentor and develop senior engineers and provide guidance on distributed systems design, debugging, and technical tradeoffs.
- Represent storage software in architecture reviews, roadmap planning, and customer-facing technical discussions.
- Design, develop, maintain, and optimize high-performance storage systems for performance, scalability, reliability, and operational simplicity.
- Implement and optimize file, block, and object storage protocol APIs, including NFS, SMB, Lustre, NVMe-oF, iSCSI, Fibre Channel, and S3.
- Develop distributed systems to manage and orchestrate storage resources across multiple solutions and redundant arrays.
- Integrate storage software with NVMe, GPU-direct storage, DPU-accelerated data paths, hardware, and system architectures.
- Troubleshoot production data center issues including performance regressions, protocol mismatches, and hardware failures.
- Contribute across requirements, system design, deployment, monitoring, and long-term maintenance.
- Build tooling for storage benchmarking, performance profiling, and capacity planning.
- Collaborate with networking, compute, control plane, Kubernetes, observability, product, and fleet engineering teams on infrastructure initiatives and data center deployments.
- Define, build, and track storage SLOs and SLIs and optimize storage for AI training checkpoint I/O, high-throughput dataset serving, and latency-sensitive inference.
- Evaluate and prototype emerging storage solutions, protocols, hardware integrations, and AI/HPC storage technologies.
Requirements
- 10+ years of storage systems engineering experience, including at least 5 years in a technical lead or Staff+ individual contributor role.
- Proven experience designing and operating storage infrastructure at scale, preferably in multi-petabyte production data center or cloud environments.
- Experience leading technical projects end to end across architecture, delivery, and cross-functional stakeholders.
- Background in high-performance computing, AI/ML infrastructure, or large-scale cloud storage.
- Strong proficiency in one or more of C, C++, Rust, or Go, with experience writing high-performance concurrent systems code and conducting code reviews.
- Experience with at least two storage protocol categories: object, block, or file, including implementing or maintaining production protocol servers or clients.
- Experience profiling and tuning storage systems for throughput, latency, and IOPS under production workloads.
- Familiarity with kernel-level storage drivers, user-space I/O frameworks, storage daemons, DPDK, and SPDK is preferred.
- Familiarity with NVMe, NVMe-oF, RDMA, DPUs, GPU-direct storage, and zero-copy data paths.
- Comfort working with physical data center infrastructure, storage arrays, cabling, rack-scale systems, and failure domains.
- Experience designing reliable storage systems, building runbooks, and driving incident response.
- Familiarity with storage observability tooling and metrics, logs, and tracing pipelines.
- Preferred experience includes NVIDIA BlueField, SuperNICs, GPUDirect Storage, Vast Data, Weka, NetApp, IBM Spectrum Scale, Ceph, CXL memory pooling, computational storage, ZNS SSDs, DAOS, Lustre, or MinIO.
- Open-source storage project contribution or maintenance experience is a plus.
Benefits
- On-site presence in San Francisco, San Jose, or Bellevue 4 days per week, with Tuesday as the designated work-from-home day.
- Generous cash and equity compensation.
- Health, dental, and vision coverage for employees and dependents.
- Wellness and commuter stipends for select roles.
- 401(k) plan with a 2% company match for USA employees.
- Flexible paid time off plan.
About Lambda
Lambda provides GPU cloud computing and on-prem AI hardware—servers, clusters, and workstations—for teams training and serving large ML models. Its products include NVIDIA H100/A100 instances, managed clusters, and the Lambda Stack software, sold via usage-based cloud pricing and hardware sales. Founded in 2012 and headquartered in San Francisco, the privately held company serves researchers, startups, enterprises, and hyperscalers.
