Shopify

Staff ML Platform Engineer - Recommendations

Shopify
Apply
3 months ago
Remote, AmericasStaff+
H1B sponsor

Responsibilities

  • Design and operate Kubernetes-based multi-node GPU training pipelines and clusters.
  • Own training reliability through checkpointing, fault tolerance, preemption recovery, and resource scheduling.
  • Optimize training performance, data-loading throughput, kernel efficiency, mixed precision, and cluster utilization.
  • Build and maintain real-time model-serving infrastructure for recommendation and LLM inference.
  • Optimize serving cost and performance through batching, model compilation, GPU right-sizing, and autoscaling.
  • Build internal ML platforms, tools, abstractions, and infrastructure standards that improve developer productivity.
  • Drive cross-team ML infrastructure strategy, write technical proposals and RFCs, mentor engineers, conduct hiring interviews, and raise technical standards.
  • Participate in shared on-call coverage for GPU cluster health, training failures, and serving availability.

Requirements

  • 7+ years of software engineering experience, including 5+ years focused on ML infrastructure or distributed systems.
  • Deep hands-on experience with GPU training at scale, distributed training, checkpointing, fault recovery, and performance tuning.
  • Strong Kubernetes experience, including pod specifications, GPU scheduling, resource quotas, scheduling-failure debugging, and stateful GPU workloads.
  • Production experience building or operating model-serving systems with real-user traffic and latency constraints.
  • Strong Python and systems fundamentals, with the ability to read and modify PyTorch training code.
  • Experience designing infrastructure abstractions used by other engineers.
  • Demonstrated technical leadership in architecture decisions, technical proposals, and engineering direction.
  • Experience mentoring engineers and raising a team's technical bar.
  • Preferred experience with SkyPilot, Ray, or similar cloud-native ML orchestration.
  • Preferred experience with vLLM, TensorRT-LLM, Triton, or equivalent LLM serving stacks.
  • Preferred production experience with model compression, including quantization, pruning, or distillation.
  • Preferred experience operating recommendation or retrieval systems at scale and building internal platforms adopted by other teams.

Benefits

  • Shared on-call rotation with end-to-end ownership of GPU cluster health, training failures, and serving availability.
  • High-trust, low-process environment with substantial ownership and autonomy.
  • Opportunity to influence ML infrastructure strategy across the organization and mentor engineers.
Shopify

About Shopify

10,000+ employees

Shopify is a leading global commerce company, providing trusted tools to start, grow, market, and manage a retail business of any size. Shopify makes commerce better for everyone with a platform and services that are engineered for reliability, while delivering a better shopping experience for consumers everywhere. Shopify powers millions of businesses in more than 175 countries and is trusted by brands such as Allbirds, Gymshark, PepsiCo, Staples, and many more. Find all our jobs here: www.shopify.com/careers

Contact me