3 hours ago
Base Salary
$300k - $405k/yr
Responsibilities
- Design and run capability, uplift, and safety evaluations for cyber-relevant model risks.
- Execute safeguard-robustness testing before major model releases.
- Analyze evaluation results and communicate findings to teams and stakeholders.
- Design, prototype, and tune detection probes for cyber misuse.
- Partner with cyber policy and engineering teams to translate policy lines into layered abuse-detection architecture and measure precision and coverage.
- Build and maintain internal tooling for running and scoring evaluations.
- Collaborate on converting evaluation findings into safeguard improvements.
Requirements
- Experience building or running evaluations, benchmarks, or test suites for software or ML systems on short, fixed timelines.
- Hands-on cybersecurity experience such as CTF participation, vulnerability research, exploit development, or security research.
- Proficiency in Python.
- Strong ability to communicate evaluation results with cross-functional and policy stakeholders.
- Bachelor’s degree in a relevant field or an equivalent combination of education, training, and experience.
- Preferred qualifications include deep offensive-security or security-research experience, AI security benchmark development, adversarial or abuse-data analysis, AI/ML evaluation frameworks, coordinated vulnerability disclosure, pre-release testing, detection-content authoring, ML-based abuse detection, and an active secret clearance or eligibility to obtain one.
Benefits
- Hybrid work with staff expected to be in an Anthropic office at least 25% of the time, with some roles requiring more office time.
- Visa sponsorship may be available, with immigration-lawyer support.
- Competitive benefits, flexible working hours, generous vacation and parental leave, optional equity donation matching, and office collaboration space.
About Anthropic
We're an AI research company that builds reliable, interpretable, and steerable AI systems. Our first product is Claude, an AI assistant for tasks at any scale. Our research interests span multiple areas including natural language, human feedback, scaling laws, reinforcement learning, code generation, and interpretability.