3 months ago
Responsibilities
- Own end-to-end model deployment pipelines, including export, validation, canary rollout, rollback, and A/B integration.
- Build CI/CD, automated testing, model validation gates, and progressive delivery for ML systems.
- Develop offline evaluation pipelines, maintain evaluation datasets, and connect offline metrics with online experimentation.
- Own training data preprocessing, feature stores, embedding tables, workflow orchestration, retraining, data quality, lineage, and freshness.
- Build monitoring and alerting for training jobs, serving endpoints, and data pipelines while defining SLOs for freshness, latency, and throughput.
- Participate in incident response, lead post-mortems, and eliminate operational toil through automation.
- Drive cross-team ML operations strategy, identify reliability risks and pipeline bottlenecks, mentor engineers, contribute to hiring, and write technical proposals and RFCs.
Requirements
- 7+ years of software engineering experience, including 5+ years focused on MLOps, data engineering, or production ML systems.
- Strong experience with ML deployment pipelines, including model export, validation, canary releases, and rollback strategies.
- Experience with ML workflow orchestration using Airflow, Dagster, Prefect, or similar tools.
- Strong Python fundamentals and familiarity with PyTorch model artifacts and training configurations.
- Production experience building monitoring, alerting, dashboards, and SLO frameworks for ML systems.
- Experience with large-scale data pipelines, batch processing, feature engineering, and data quality validation.
- Working proficiency with Kubernetes, including debugging pod failures, understanding resource scheduling, and navigating GPU workloads.
- Demonstrated technical leadership through operational strategy, technical proposals, engineering influence, mentoring, and raising reliability standards.
- Preferred experience with BigQuery or equivalent data warehouses, experiment tracking, A/B testing frameworks, recommendation or retrieval systems, model compression, and cloud-native ML orchestration such as SkyPilot or Ray.
Benefits
- Shared on-call rotation covering pipeline failures, data freshness alerts, deployment rollbacks, and evaluation integrity.
- High-ownership environment with autonomy, direct pairing with ML Engineers, and a preference for automation over runbooks.
Tech Stack
Categories
Data EngineeringML Engineering
About Shopify
Shopify is a leading global commerce company, providing trusted tools to start, grow, market, and manage a retail business of any size. Shopify makes commerce better for everyone with a platform and services that are engineered for reliability, while delivering a better shopping experience for consumers everywhere. Shopify powers millions of businesses in more than 175 countries and is trusted by brands such as Allbirds, Gymshark, PepsiCo, Staples, and many more. Find all our jobs here: www.shopify.com/careers
