3 days ago
Base Salary
$144k - $194k/yr
Responsibilities
- Design, develop, and optimize machine learning models and frameworks for deployment on custom ML hardware accelerators.
- Build distributed inference support for PyTorch in the AWS Neuron SDK and tune Llama, DeepSeek, and other LLM families.
- Design and implement high-performance kernels and ML operation features using Neuron architecture and programming models.
- Participate in architecture, implementation, profiling, hardware-specific optimization, testing, and production deployment across the ML system lifecycle.
- Build infrastructure to analyze and onboard models with diverse architectures.
- Optimize system performance, memory usage, latency, throughput, and efficiency across multiple generations of Neuron hardware.
- Implement optimizations including fusion, sharding, tiling, and scheduling.
- Conduct unit and end-to-end model testing and support continuous deployment and release pipelines.
- Debug software and performance issues, identify bottlenecks, and resolve root causes of defects.
- Work directly with customers to enable and optimize machine learning models on AWS accelerators.
- Collaborate with compiler, runtime, framework, hardware, applied science, systems engineering, and product management teams.
Requirements
- At least 3 years of non-internship professional software development experience.
- At least 3 years of experience designing or architecting new and existing systems, including design patterns, reliability, and scaling.
- Bachelor's degree in computer science or equivalent.
- Fundamental knowledge of machine learning and LLM architectures, training and inference lifecycles, and model execution optimization.
- Software development experience in C++ or Python, with experience in at least one required.
- Strong understanding of system performance, memory management, and parallel computing principles.
- Proficiency in debugging, profiling, and applying software engineering practices in large-scale systems.
- Preferred familiarity with PyTorch, JIT compilation, and AOT tracing.
- Preferred familiarity with CUDA kernels or equivalent ML or low-level kernels.
- Preferred experience developing performant kernels with technologies such as CUTLASS or FlashInfer.
- Preferred familiarity with syntax and tile-level semantics similar to Triton.
- Preferred production experience with online or offline inference serving using vLLM, SGLang, TensorRT, or similar platforms.
- Preferred deep understanding of computer architecture, operating-system-level software, and parallel computing.
Benefits
- The role includes sign-on payments and restricted stock units in addition to base salary.
- Benefits include medical, dental, vision, prescription, life and AD&D insurance, supplemental life plan options, an employee assistance program, mental health support, a medical advice line, flexible spending accounts, adoption and surrogacy reimbursement, 401(k) matching, paid time off, and parental leave.
- The position is based in Seattle, Washington, United States.