Amazon

Sr. SDE, Edge AI ML Platform, Edge AI and Science

Amazon
Apply
1 day ago
Vancouver, CanadaSenior
H1B Sponsor

Responsibilities

  • Lead the design and delivery of distributed ML platform services and libraries for model ingestion, optimization, training, evaluation, packaging, and deployment.
  • Define stable APIs and architecture boundaries for integrating research algorithms with training, infrastructure, and deployment systems.
  • Design and scale distributed training across data, tensor, pipeline, and model parallelism on multi-node GPU clusters.
  • Improve training throughput, GPU utilization, memory efficiency, communication performance, failure recovery, and developer iteration speed.
  • Connect distributed training with distillation, quantization, pruning, and other model optimization techniques.
  • Build evaluation, artifact, deployment, automated validation, regression testing, observability, and release workflows for GPU-intensive workloads.
  • Profile and optimize end-to-end system performance with applied scientists and GPU kernel engineers.
  • Establish production operational mechanisms including metrics, alarms, runbooks, on-call practices, and root-cause correction.
  • Partner with model, compiler, runtime, hardware, security, infrastructure, and product teams to manage dependencies and deliver multi-team programs.
  • Lead architecture discussions, write technical designs, evaluate trade-offs, build consensus, mentor engineers, and support engineering-team recruiting and development.

Requirements

  • At least 5 years of non-internship professional software development experience.
  • At least 5 years of programming experience with at least one software programming language.
  • At least 5 years of experience leading the design or architecture of new and existing systems, including design patterns, reliability, and scaling.
  • Experience as a mentor, tech lead, or engineering-team leader.
  • Experience designing or building distributed systems or high-performance computing systems.
  • Preferred: at least 5 years of full software development lifecycle experience, including coding standards, code reviews, source control management, build processes, testing, and operations.
  • Preferred: experience building distributed ML training, inference, evaluation, or data platforms using PyTorch, TensorFlow, JAX, NeMo, or Megatron.
  • Preferred: experience with containers, Kubernetes, AWS infrastructure, observability, and production operations.
  • Preferred: experience with model compression, quantization, knowledge distillation, model compilation, or edge deployment.
  • Preferred: experience designing extensible platform APIs and delivering systems with science, hardware, compiler, or product teams.

Benefits

  • Comprehensive health insurance including medical, dental, vision, prescription, basic life, and AD&D insurance.
  • Registered Retirement Savings Plan (RRSP) and Deferred Profit Sharing Plan (DPSP).
  • Paid time off and other health and well-being resources.
  • The position is based in Vancouver, British Columbia, Canada.

Tech Stack

Amazon

About Amazon

10,000+ employees
Contact me