2 months ago
Base Salary
$226k - $355k/yr
Responsibilities
- Partner with account executives to lead complex enterprise and digital-native customer deals and build relationships with technical leaders.
- Evaluate customer architectures, identify bottlenecks, and design end-to-end GPU cloud solutions.
- Author proposals and architecture diagrams and collaborate on bills of materials and rack elevations for multi-node GPU clusters.
- Lead hands-on technical proof-of-concepts, custom prototypes, and benchmark evaluations for training and inference workloads.
- Architect and optimize AI/ML workloads, including data ingestion, distributed training, inference optimization, and observability.
- Advise customers on high-performance networking, distributed storage, and cluster topologies to maximize GPU utilization.
- Represent the technical voice of the customer to Lambda’s product and engineering teams.
- Create technical enablement materials, whitepapers, architectural blueprints, and workshops.
- Represent Lambda at conferences, webinars, and technical community events.
Requirements
- 8+ years of experience designing, deploying, and scaling enterprise cloud infrastructure.
- 4+ years of experience in a Solution Architect, Solution Engineer, or technical customer-facing capacity supporting complex cloud environments.
- 3+ years of hands-on experience architecting and deploying cloud-based AI/ML workloads.
- Experience deploying, benchmarking, and optimizing workloads on NVIDIA GPU architectures such as HGX platforms and NVLink using PyTorch, NeMo, vLLM, and TensorRT-LLM.
- Strong experience with Kubernetes, Docker, SLURM, Terraform, and Ansible.
- Deep knowledge of high-speed networking, distributed file systems, security, and cost optimization, including InfiniBand, RoCE, NFS, NVMe-oF, Weka, and VAST.
- Experience coding in Python, Go, C, C++, CUDA, or a similar programming language.
- Experience partnering with account executives on complex cloud deals and presenting technical architectures to C-level stakeholders.
- Demonstrated organizational or multi-departmental impact and experience mentoring junior solution engineers or architects.
- Nice-to-have experience with LLM fine-tuning, algorithm selection, pipeline design, distributed training, 3D parallelism, or Megatron-LM.
- Nice-to-have experience with product launches, go-to-market initiatives, technical whitepapers, benchmarks, RESTful APIs, gRPC, and service-oriented cloud architectures.
Benefits
- Requires working from the San Francisco, San Jose, or Bellevue office 4 days per week, with Tuesday designated as the work-from-home day.
- Health, dental, and vision coverage for employees and dependents.
- Wellness and commuter stipends for select roles.
- 401(k) plan with a 2% company match for U.S. employees.
- Flexible paid time off plan.
- Generous cash and equity compensation.
About Lambda
Lambda provides GPU cloud computing and on-prem AI hardware—servers, clusters, and workstations—for teams training and serving large ML models. Its products include NVIDIA H100/A100 instances, managed clusters, and the Lambda Stack software, sold via usage-based cloud pricing and hardware sales. Founded in 2012 and headquartered in San Francisco, the privately held company serves researchers, startups, enterprises, and hyperscalers.
