4 hours ago
Bellevue, WA, USAMid Level / Senior
H1B sponsor
Base Salary
$200k - $288k/yr
Responsibilities
- Design and build the public training APIs, SDK, control plane, and GPU data plane.
- Scale multi-tenant scheduling, placement, and capacity-aware routing across regional GPU pools with built-in fault tolerance.
- Optimize training, inference, and reinforcement-learning loops for responsiveness, throughput, GPU utilization, and cost efficiency.
- Productionize research techniques into reliable, composable training and inference components for enterprise customers.
Requirements
- 3+ years of experience building and shipping production ML systems for the Intermediate level or 6+ years for the Senior level.
- Strong distributed-systems and infrastructure experience designing scalable, fault-tolerant services and operating them on Kubernetes in production.
- Familiarity with GPU and LLM infrastructure including PyTorch, DeepSpeed/FSDP, Ray, CUDA/NCCL, and vLLM, with the ability to debug across data, infrastructure, and GPU layers.
- Demonstrated ability to harden complex systems for reliability, throughput, and cost efficiency.
- Bachelor’s degree in computer science or a related field; a master’s degree or PhD is a plus.
- Hands-on LLM post-training or modeling experience is a bonus.
Tech Stack
Categories
About Snowflake
Snowflake builds a cloud-native data platform used by enterprises to store, integrate, share, and analyze data across AWS, Azure, and Google Cloud. Its core products span data warehousing, data lakes, data engineering, and governed data sharing, sold via consumption-based subscriptions. Founded in 2012 and publicly traded on the NYSE (SNOW) following a 2020 IPO, Snowflake supports analytics and data application workloads across industries.
