9 months ago
Base Salary
$300k - $405k/yr
Responsibilities
- Design and implement reinforcement learning environments
- Develop novel approaches for safe AI models and realize them in code
- Conduct experiments and evaluations
- Deliver research and engineering work into production training runs
- Advance model capabilities in secure coding, vulnerability remediation, and defensive cybersecurity
- Collaborate with researchers, engineers, and cybersecurity specialists inside and outside Anthropic
Requirements
- Experience in cybersecurity research
- Experience with machine learning
- Strong software engineering skills
- Ability to balance research exploration with engineering implementation
- Passion for AI’s potential and commitment to safe, beneficial AI systems
- At least a bachelor’s degree in a related field or equivalent experience
- Professional experience in security engineering, fuzzing, detection and response, or other applied defensive work is a strong plus
- Experience participating in or building CTF competitions and cyber ranges is a strong plus
- Academic research experience in cybersecurity is a strong plus
- Familiarity with reinforcement learning techniques and environments is a strong plus
- Familiarity with LLM training methodologies is a strong plus
Benefits
- Annual base compensation of $300,000–$405,000 USD, excluding equity and incentive compensation
- Hybrid policy requiring staff to work from an office at least 25% of the time, with some roles requiring more
- Visa sponsorship may be available, supported by an immigration lawyer
- Optional equity donation matching
- Generous vacation and parental leave
- Flexible working hours
- Collaborative office space
Categories
About Anthropic
We're an AI research company that builds reliable, interpretable, and steerable AI systems. Our first product is Claude, an AI assistant for tasks at any scale. Our research interests span multiple areas including natural language, human feedback, scaling laws, reinforcement learning, code generation, and interpretability.