7 months ago
Base Salary
$320k - $485k/yr
Responsibilities
- Design and build evaluation systems measuring model capabilities across diverse coding tasks
- Build tooling and infrastructure for researchers to run experiments at scale
- Develop pipelines for data collection, processing, and analysis
- Create internal tools that improve researcher productivity and accelerate iteration
- Bridge product and research by using product intuition to identify important capabilities
- Translate research questions into engineering solutions with researchers
- Own systems end-to-end from design through production reliability
Requirements
- At least 5 years of work experience
- Experience building and owning complex systems such as pipelines, infrastructure, or software that orchestrates multiple components and handles significant state and logic
- Ability to work independently in fast-paced, high-intensity environments and drive ambiguous problems to completion
- Strong interest in agentic coding tools and intuition about model capabilities and limitations
- Care for correctness and reliability in systems
- Bachelor's degree in a related field or equivalent experience
- Strong candidates may have experience with evaluation frameworks, reinforcement learning systems, research computing or scientific infrastructure, quantitative fields, Python, or TypeScript
Benefits
- Hybrid work with at least 25% of time expected in an office
- Visa sponsorship with immigration-lawyer support
- Competitive compensation and benefits
- Optional equity donation matching
- Generous vacation and parental leave
- Flexible working hours
- Collaborative office space
Tech Stack
About Anthropic
We're an AI research company that builds reliable, interpretable, and steerable AI systems. Our first product is Claude, an AI assistant for tasks at any scale. Our research interests span multiple areas including natural language, human feedback, scaling laws, reinforcement learning, code generation, and interpretability.