1 day ago
Responsibilities
- Lead the design and delivery of distributed ML platform services and libraries for model ingestion, optimization, training, evaluation, packaging, and deployment.
- Define stable APIs and architecture boundaries for integrating research algorithms with training, infrastructure, and deployment systems.
- Design and scale distributed training across data, tensor, pipeline, and model parallelism on multi-node GPU clusters.
- Improve training throughput, GPU utilization, memory efficiency, communication performance, failure recovery, and developer iteration speed.
- Connect distributed training with distillation, quantization, pruning, and other model optimization techniques.
- Build evaluation, artifact, deployment, automated validation, regression testing, observability, and release workflows for GPU-intensive workloads.
- Profile and optimize end-to-end system performance with applied scientists and GPU kernel engineers.
- Establish production operational mechanisms including metrics, alarms, runbooks, on-call practices, and root-cause correction.
- Partner with model, compiler, runtime, hardware, security, infrastructure, and product teams to manage dependencies and deliver multi-team programs.
- Lead architecture discussions, write technical designs, evaluate trade-offs, build consensus, mentor engineers, and support engineering-team recruiting and development.
Requirements
- At least 5 years of non-internship professional software development experience.
- At least 5 years of programming experience with at least one software programming language.
- At least 5 years of experience leading the design or architecture of new and existing systems, including design patterns, reliability, and scaling.
- Experience as a mentor, tech lead, or engineering-team leader.
- Experience designing or building distributed systems or high-performance computing systems.
- Preferred: at least 5 years of full software development lifecycle experience, including coding standards, code reviews, source control management, build processes, testing, and operations.
- Preferred: experience building distributed ML training, inference, evaluation, or data platforms using PyTorch, TensorFlow, JAX, NeMo, or Megatron.
- Preferred: experience with containers, Kubernetes, AWS infrastructure, observability, and production operations.
- Preferred: experience with model compression, quantization, knowledge distillation, model compilation, or edge deployment.
- Preferred: experience designing extensible platform APIs and delivering systems with science, hardware, compiler, or product teams.
Benefits
- Comprehensive health insurance including medical, dental, vision, prescription, basic life, and AD&D insurance.
- Registered Retirement Savings Plan (RRSP) and Deferred Profit Sharing Plan (DPSP).
- Paid time off and other health and well-being resources.
- The position is based in Vancouver, British Columbia, Canada.
