13 hours ago
Base Salary
$168k - $227k/yr
Responsibilities
- Build distributed inference support for PyTorch in the Neuron SDK.
- Design, develop, and optimize machine learning models, including GPT, Kimi, and Qwen, on custom AI accelerators.
- Participate in distributed system architecture, implementation, performance profiling, low-level optimization, and production deployment.
- Build infrastructure to analyze and onboard models with diverse architectures.
- Design and implement high-performance kernels and ML operation features using Neuron programming models.
- Analyze and optimize system performance, memory usage, and bottlenecks across multiple generations of Neuron hardware.
- Implement fusion, sharding, tiling, scheduling, graph transformations, and other optimization techniques.
- Work with customers to enable and optimize their ML models on AWS accelerators.
- Collaborate across software, hardware, open-source, and customer teams; participate in design discussions and code reviews.
- Create metrics, implement automation, debug software defects, and mentor engineers on optimization.
Requirements
- Bachelor's degree.
- 5+ years of non-internship professional software development experience.
- Knowledge of Python and/or C++ programming.
- 5+ years of experience leading the design or architecture of new and existing systems.
- Experience debugging, profiling, and applying software engineering best practices in large-scale systems.
- Knowledge of system performance, memory management, and parallel computing principles.
- Experience owning a performance optimization roadmap and mentoring engineers on optimization.
- Preferred: master's degree in computer science or equivalent.
- Preferred: knowledge of machine learning model architecture, inference, ML and LLM fundamentals, transformer architecture, and training/inference lifecycles.
- Preferred: hands-on PyTorch development.
- Preferred: experience scaling workloads across multi-GPU and multi-node topologies using NCCL and tensor, pipeline, or expert parallelism.
- Preferred: experience writing and optimizing custom CUDA/Triton kernels for tensor operations.
Benefits
- Comprehensive health insurance, including medical, dental, vision, prescription, life, and AD&D coverage.
- 401(k) matching, paid time off, parental leave, and adoption and surrogacy reimbursement coverage.
- Employee assistance, mental health support, medical advice line, and flexible spending accounts.
- The role is based in Seattle, Washington, USA.
- Compensation also includes sign-on payments and restricted stock units, with benefits provided by Amazon.
