29 days ago
Base Salary
$266k - $395k/yr
Responsibilities
- Design and build a vendor-agnostic control plane that provisions, scales, heals, and meters storage across multiple storage platforms.
- Create an abstraction layer for vendor-specific APIs, failure semantics, quality-of-service controls, and telemetry formats.
- Develop reconciliation loops, Kubernetes controllers, operators, and custom schedulers for capacity, tenancy, encryption domains, and placement.
- Implement end-to-end multi-tenant isolation, including namespace and subsystem partitioning, tenant quality of service, rate limiting, credential and key lifecycle management, blast-radius containment, and noisy-neighbor detection.
- Design a PCIe-topology-aware, NUMA-aware, and failure-domain-aware capacity and placement engine.
- Define service-level indicators and objectives, detect fleet-wide performance regressions, and build observability for petabyte-scale storage infrastructure.
Requirements
- Bachelor's or master's degree in computer science or a related field.
- 5+ years of experience in software development for storage systems.
- Proven experience with distributed systems programming, including load balancing, data durability, consensus algorithms, fault tolerance, and data consistency.
- Strong programming skills in C, C++, Go, or Python.
- Experience with Linux kernel internals and system-level programming.
- Experience with storage protocols such as S3 or NFS and file systems such as Ceph, DAOS, or similar.
- Familiarity with Docker, Kubernetes, and production workloads in containerized environments.
- Familiarity with CI/CD and QA practices for distributed systems development.
- Experience with AI/ML workloads, data center networking, InfiniBand, RoCE, storage performance tuning, GPUs, DPUs, or named storage platforms is preferred.
- Experience with VAST Data, WEKA, DDN, Pure, NetApp, IBM Storage Scale, large-scale Ceph, CXL memory pooling, computational storage, ZNS SSDs, EDSFF, or relevant systems research publications is a plus.
Benefits
- Health, dental, and vision coverage for employees and dependents.
- Wellness and commuter stipends for select roles.
- 401(k) plan with a 2% company match for USA employees.
- Flexible paid time off plan.
- Generous cash and equity compensation.
- Hybrid schedule requiring presence in the San Francisco, San Jose, or Bellevue office 4 days per week, with Tuesday designated as the work-from-home day.
About Lambda
Lambda provides GPU cloud computing and on-prem AI hardware—servers, clusters, and workstations—for teams training and serving large ML models. Its products include NVIDIA H100/A100 instances, managed clusters, and the Lambda Stack software, sold via usage-based cloud pricing and hardware sales. Founded in 2012 and headquartered in San Francisco, the privately held company serves researchers, startups, enterprises, and hyperscalers.
