over 1 year ago
Base Salary
$315k - $560k/yr
Responsibilities
- Identify novel systems problems arising from running ML algorithms at scale
- Develop systems that optimize throughput and robustness of large distributed systems
- Implement low-latency, high-throughput sampling for large language models
- Implement GPU kernels for low-precision model inference
- Develop load-balancing algorithms to improve serving efficiency
- Build quantitative models of system performance
- Design and implement fault-tolerant distributed systems with complex network topologies
- Debug kernel-level network latency spikes in containerized environments
Requirements
- Significant software engineering or machine learning experience, particularly at supercomputing scale
- Experience with high-performance, large-scale ML systems, GPU or accelerator programming, ML framework internals, OS internals, or transformer-based language modeling is valuable
- Interest in learning more about machine learning research
- Strong collaboration and communication skills
Benefits
- Hybrid work policy requiring staff to be in an office at least 25% of the time
- Visa sponsorship may be available, with immigration lawyer support
- Competitive benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and office collaboration space
Categories
About Anthropic
We're an AI research company that builds reliable, interpretable, and steerable AI systems. Our first product is Claude, an AI assistant for tasks at any scale. Our research interests span multiple areas including natural language, human feedback, scaling laws, reinforcement learning, code generation, and interpretability.