Amazon

Software Development Engineer, AI/ML, AWS Neuron, Model Inference

Amazon
Apply
7 days ago
Cupertino, CA, USA or Seattle, WA, USAMid Level
H1B sponsor

Base Salary

$165k - $224k/yr

Responsibilities

  • Design, develop, and optimize machine learning models and frameworks for deployment on custom ML hardware accelerators.
  • Build distributed inference support and systematically analyze and onboard models with diverse architectures.
  • Design and implement high-performance kernels and ML operation features using the Neuron architecture and programming models.
  • Analyze and optimize system performance, memory usage, latency, throughput, and efficiency across multiple generations of Neuron hardware.
  • Conduct profiling, debugging, unit testing, end-to-end model testing, production deployment, and continuous releases.
  • Implement optimization techniques including fusion, sharding, tiling, and scheduling.
  • Work directly with customers to enable and optimize their ML models on AWS accelerators.
  • Collaborate with compiler, runtime, framework, hardware, applied science, and product teams on inference capabilities and optimization techniques.
  • Create metrics and automation, resolve software defects, participate in design discussions and code reviews, and mentor engineers.

Requirements

  • Bachelor's degree in computer science or equivalent.
  • At least 3 years of non-internship professional software development experience.
  • At least 3 years of non-internship experience designing or architecting new and existing systems, including design patterns, reliability, and scaling.
  • Fundamental knowledge of machine learning and LLM architectures, training and inference lifecycles, and model execution optimization.
  • Software development experience in C++ or Python, with at least one language required.
  • Strong understanding of system performance, memory management, and parallel computing principles.
  • Deep understanding of computer architecture, operating-systems-level software, and parallel computing.
  • Proficiency in debugging, profiling, and software engineering practices for large-scale systems.
  • Preferred: Master's degree or Ph.D. in computer science or equivalent.
  • Preferred: familiarity with PyTorch, JIT compilation, AOT tracing, CUDA kernels or equivalent low-level kernels, CUTLASS, FlashInfer, Triton-like syntax and tile-level semantics, and production inference serving with vLLM, SGLang, TensorRT, or similar platforms.

Benefits

  • Amazon offers health insurance, 401(k) matching, paid time off, parental leave, sign-on payments, and restricted stock units.
  • Benefits include medical, dental, vision, prescription, life and AD&D insurance options, an employee assistance program, mental health support, a medical advice line, flexible spending accounts, and adoption and surrogacy reimbursement coverage.
  • The role is based in Cupertino, California, United States; no remote or hybrid arrangement is stated.

Categories

Amazon

About Amazon

10,000+ employees

Amazon builds and operates a global e-commerce marketplace, logistics network, and consumer devices, and provides cloud computing via AWS for businesses and developers. The company earns revenue from online retail, third‑party seller services, subscriptions like Prime, advertising, and AWS usage. Founded in 1994 and headquartered in Seattle, it is publicly traded on NASDAQ (AMZN) and serves customers in dozens of countries.

Contact me