11 months ago
Base Salary
$320k - $425k/yr
Responsibilities
- Develop monitoring systems to detect unwanted behaviors from API partners and potentially take automated enforcement actions.
- Surface detected behaviors in internal dashboards for analyst review.
- Build abuse detection mechanisms and infrastructure.
- Surface abuse patterns to research teams to support model hardening during training.
- Build robust, reliable, multi-layered defenses for real-time improvement of safety mechanisms at scale.
- Analyze user reports of inappropriate content or accounts.
Requirements
- Bachelor’s degree in Computer Science, Software Engineering, or a comparable field, or equivalent experience.
- 5–10+ years of experience in a software engineering position, preferably focused on integrity, spam, fraud, or abuse detection and mitigation.
- Proficiency in Python and TypeScript.
- Ability to work across the stack.
- Strong communication skills and ability to explain complex technical concepts to non-technical stakeholders.
Benefits
- Full-time employees receive equity, benefits, and may receive incentive compensation.
- Optional equity donation matching is available.
- Generous vacation and parental leave are offered.
- Flexible working hours are available.
- The role follows a location-based hybrid policy requiring staff to be in an office at least 25% of the time, with some roles requiring more.
- Visa sponsorship may be available, supported by an immigration lawyer.
- The company provides office space for collaboration.
Tech Stack
Categories
About Anthropic
We're an AI research company that builds reliable, interpretable, and steerable AI systems. Our first product is Claude, an AI assistant for tasks at any scale. Our research interests span multiple areas including natural language, human feedback, scaling laws, reinforcement learning, code generation, and interpretability.