over 1 year ago
Base Salary
$315k - $560k/yr
Responsibilities
- Identify and address performance issues across research, training, and inference ML systems
- Design and optimize TPU kernels
- Provide researchers with feedback on how model changes affect performance
- Implement low-latency, high-throughput sampling for large language models
- Adapt models for low-precision inference
- Build quantitative models of system performance
- Design and implement custom collective communication algorithms
- Debug kernel performance at the assembly level
Requirements
- Significant experience optimizing machine-learning systems for TPUs, GPUs, or other accelerators
- Experience with high-performance, large-scale ML systems
- Experience designing and implementing kernels for TPUs or other ML accelerators
- Deep understanding of accelerators, potentially including computer architecture
- Experience with ML framework internals and transformer-based language modeling is valuable
- Bachelor’s degree in a related field or equivalent experience
- Strong systems problem-solving and low-level optimization experience
Benefits
- Hybrid work with staff expected in an office at least 25% of the time
- Visa sponsorship may be available, with immigration lawyer support
- Competitive compensation and benefits
- Optional equity donation matching
- Generous vacation and parental leave
- Flexible working hours
- Collaborative office environment
Tech Stack
Assembly
Categories
About Anthropic
We're an AI research company that builds reliable, interpretable, and steerable AI systems. Our first product is Claude, an AI assistant for tasks at any scale. Our research interests span multiple areas including natural language, human feedback, scaling laws, reinforcement learning, code generation, and interpretability.