3 months ago
Base Salary
$160k - $230k/yr
Responsibilities
- Design and build across the full stack, including public training APIs and SDKs, the control plane, and the GPU data plane.
- Scale multi-tenant scheduling, placement, and capacity-aware routing across regional GPU pools with built-in fault tolerance.
- Optimize training, inference, and reinforcement-learning loops and maintain responsive data-plane performance under heavy concurrent load.
- Partner with Snowflake Research to productionize state-of-the-art training and inference techniques into reliable, composable components for enterprise-scale customers.
Requirements
- At least five years of experience building and shipping production ML systems.
- Strong distributed systems and infrastructure foundation, including designing scalable, fault-tolerant services and operating them on Kubernetes in production.
- Familiarity with GPU and LLM infrastructure including PyTorch, DeepSpeed/FSDP, Ray, CUDA/NCCL, and vLLM, with the ability to debug across data, infrastructure, and GPU layers.
- Demonstrated ability to harden complex systems for reliability, throughput, and cost efficiency.
- Bachelor’s degree in Computer Science or a related field.
- Hands-on LLM post-training or modeling experience is a bonus.
- A master’s degree or PhD is a plus.
Tech Stack
Categories
About Snowflake
Snowflake builds a cloud-native data platform used by enterprises to store, integrate, share, and analyze data across AWS, Azure, and Google Cloud. Its core products span data warehousing, data lakes, data engineering, and governed data sharing, sold via consumption-based subscriptions. Founded in 2012 and publicly traded on the NYSE (SNOW) following a 2020 IPO, Snowflake supports analytics and data application workloads across industries.
