17 hours ago
Base Salary
$193k - $262k/yr
Responsibilities
- Build distributed inference support for PyTorch in the AWS Neuron SDK.
- Develop, optimize, and deploy machine learning models and frameworks on Trainium and Inferentia accelerators.
- Design high-performance kernels and apply fusion, sharding, tiling, and scheduling optimizations.
- Analyze system performance, memory usage, and bottlenecks across multiple Neuron hardware generations using profiling tools.
- Build infrastructure to analyze and onboard models with diverse architectures.
- Perform unit and end-to-end testing and support continuous deployment and releases.
- Work directly with customers to enable and optimize their machine learning models.
- Collaborate across compiler, runtime, framework, hardware, applied science, and product teams.
- Mentor engineers and contribute to architecture, code reviews, automation, and root-cause analysis.
Requirements
- Bachelor's degree in computer science or equivalent.
- At least 5 years of non-internship professional software development experience.
- At least 5 years of programming experience with a software programming language.
- At least 5 years of experience leading the design or architecture of new and existing systems.
- Experience mentoring, serving as a technical lead, or leading an engineering team.
- Machine learning and LLM fundamentals, including model architecture, training, inference lifecycles, and execution optimization experience.
- Software development experience in C++ or Python.
- Strong understanding of system performance, memory management, parallel computing, computer architecture, and operating-system-level software.
- Proficiency in debugging, profiling, and software engineering practices for large-scale systems.
- Preferred: master's degree in computer science or equivalent.
- Preferred: at least 5 years of full software development lifecycle experience, including coding standards, code reviews, source control, build processes, testing, and operations.
- Preferred: familiarity with PyTorch, JIT compilation, AOT tracing, CUDA kernels, or equivalent low-level kernels.
- Preferred: performant kernel development experience with technologies such as CUTLASS, FlashInfer, or Triton-like systems.
- Preferred: production experience with online or offline inference serving using vLLM, SGLang, TensorRT, or similar platforms.
Benefits
- Comprehensive medical, dental, vision, prescription, life, AD&D, employee assistance, mental health, medical advice line, and flexible spending benefits.
- 401(k) matching, paid time off, parental leave, and adoption and surrogacy reimbursement coverage.
- Sign-on payments and restricted stock units are included in the Amazon compensation package.
- Role is based in Cupertino, California.
Categories
About Amazon
Amazon builds and operates a global e-commerce marketplace, logistics network, and consumer devices, and provides cloud computing via AWS for businesses and developers. The company earns revenue from online retail, third‑party seller services, subscriptions like Prime, advertising, and AWS usage. Founded in 1994 and headquartered in Seattle, it is publicly traded on NASDAQ (AMZN) and serves customers in dozens of countries.
