Amazon

Software Development Engineer, AI/ML, AWS Neuron, Model Inference

Amazon
Apply
17 hours ago
Cupertino, CA, USAMid Level
H1B sponsor

Base Salary

$165k - $224k/yr

Responsibilities

  • Design, develop, and optimize machine learning models and frameworks for deployment on custom ML accelerators.
  • Build distributed inference support and systematically analyze and onboard diverse model architectures.
  • Design and implement high-performance kernels and ML operations using Neuron programming models.
  • Profile and optimize system performance, memory usage, latency, throughput, and hardware efficiency across Neuron generations.
  • Implement optimizations including fusion, sharding, tiling, and scheduling.
  • Perform unit and end-to-end model testing and support continuous deployment and release pipelines.
  • Debug software defects and performance bottlenecks and develop automation and performance metrics.
  • Work with compiler, runtime, framework, hardware, applied science, and product teams.
  • Work directly with customers to enable and optimize machine learning models on AWS accelerators.
  • Participate in architecture and design discussions, code reviews, and technical communication with internal and external stakeholders.
  • Mentor experienced engineers and contribute technical input to system and business decisions.

Requirements

  • Bachelor's degree in computer science or equivalent.
  • At least 3 years of non-internship professional software development experience.
  • At least 3 years of non-internship experience designing or architecting new and existing systems.
  • Fundamentals of machine learning and LLM architectures, training, inference lifecycles, and model-execution optimization experience.
  • Software development experience in C++ or Python.
  • Strong understanding of system performance, memory management, and parallel computing principles.
  • Deep understanding of computer architecture, operating-system-level software, and parallel computing.
  • Proficiency in debugging, profiling, and software engineering practices for large-scale systems.
  • Preferred master's degree or Ph.D. in computer science or equivalent.
  • Preferred familiarity with PyTorch, JIT compilation, AOT tracing, CUDA kernels, or equivalent low-level ML kernels.
  • Preferred experience with performant kernel development using technologies such as CUTLASS, FlashInfer, or Triton-like tile-level semantics.
  • Preferred production experience with online or offline inference serving using vLLM, SGLang, TensorRT, or similar platforms.

Benefits

  • The role is based in Cupertino, California.
  • Amazon offers health insurance including medical, dental, vision, prescription, life and AD&D insurance, employee assistance, mental health support, a medical advice line, flexible spending accounts, and adoption and surrogacy reimbursement.
  • Benefits include 401(k) matching, paid time off, and parental leave.
  • The compensation package may include sign-on payments and restricted stock units.
Amazon

About Amazon

10,000+ employees

Amazon builds and operates a global e-commerce marketplace, logistics network, and consumer devices, and provides cloud computing via AWS for businesses and developers. The company earns revenue from online retail, third‑party seller services, subscriptions like Prime, advertising, and AWS usage. Founded in 1994 and headquartered in Seattle, it is publicly traded on NASDAQ (AMZN) and serves customers in dozens of countries.

Contact me