3 months ago
Base Salary
$266k - $395k/yr
Responsibilities
- Design, develop, and maintain high-performance, scalable, and reliable storage software.
- Implement and optimize file, block, and object storage protocol APIs.
- Develop distributed systems for managing and orchestrating storage resources across multiple solutions and redundant arrays.
- Integrate storage software with NVMe and GPU-direct storage alongside hardware and systems architects.
- Troubleshoot and debug complex issues in production data-center environments.
- Contribute across the full software development lifecycle from requirements and design through deployment and maintenance.
- Coordinate with storage software, networking, control plane, MK8s, observability, compute, fleet engineering, and product teams.
- Deploy and maintain distributed storage solutions for AI and machine-learning workloads.
- Build and track SLOs and SLIs with the observability team.
- Research and optimize storage protocols for AI inference, training, and scientific-computing applications.
- Provide technical leadership and support cross-functional infrastructure initiatives and new data-center deployments.
Requirements
- 10+ years of experience in storage engineering, including at least 5+ years in a management or lead role.
- Experience serving object, block, or file storage protocols such as S3, iSCSI, NFS, SMB, or Lustre.
- Professional individual-contributor experience as a storage engineer or storage SRE.
- Experience with systems-level programming, storage protocol APIs, storage performance optimization, physical infrastructure, and operational responsibilities.
- Familiarity with NVMe, RDMA, and DPUs and their role in optimizing storage performance.
- Experience building high-performance teams through hiring, upskilling, skills redundancy, performance management, and expectation setting.
- Preferred experience driving cross-functional engineering management initiatives and large projects.
- Preferred experience with NVIDIA SuperNIC DPUs, edge caching, or GPUDirect Storage.
- Preferred deep experience with Vast, Weka, NetApp, or Ceph in HPC or AI infrastructure environments.
- Preferred experience implementing Ceph at a scale greater than 100PB.
- Preferred experience driving organizational improvements or training and managing managers.
Benefits
- Requires presence in the San Francisco office 4 days per week, with Tuesday as the designated work-from-home day.
- Health, dental, and vision coverage for employees and dependents.
- Wellness and commuter stipends for select roles.
- 401(k) plan with a 2% company match for USA employees.
- Flexible paid time off plan.
- Cash and equity compensation are offered, with no specific amounts stated.
About Lambda
Lambda provides GPU cloud computing and on-prem AI hardware—servers, clusters, and workstations—for teams training and serving large ML models. Its products include NVIDIA H100/A100 instances, managed clusters, and the Lambda Stack software, sold via usage-based cloud pricing and hardware sales. Founded in 2012 and headquartered in San Francisco, the privately held company serves researchers, startups, enterprises, and hyperscalers.
