12 hours ago
Base Salary
$320k - $485k/yr
Responsibilities
- Develop monitoring systems to detect unwanted behaviors from API partners and support automated enforcement actions.
- Build internal dashboards that surface detected issues for analyst review.
- Build abuse-detection mechanisms and supporting infrastructure.
- Surface abuse patterns to research teams to help harden models during training.
- Build reliable, multilayered safety defenses that improve in real time and operate at scale.
Requirements
- Bachelor’s degree in Computer Science, Software Engineering, or a comparable qualification, or equivalent experience.
- Relevant field of study demonstrated through coursework, training, or professional experience.
- Proficiency in Python and TypeScript.
- Ability to work across the stack.
- Strong communication skills and ability to explain complex technical concepts to non-technical stakeholders.
- Strong candidates may have 8+ years of software engineering experience.
- Preferred experience includes integrity, spam, fraud, abuse detection and mitigation, trust and safety mechanisms for AI/ML systems, prompt engineering, jailbreak attacks, adversarial inputs, and custom internal tooling for operational teams.
Benefits
- Annual base salary range of $320,000–$485,000 USD.
- Visa sponsorship is offered, with reasonable efforts made after an offer and support from an immigration lawyer.
- Hybrid work policy requiring staff to be in an Anthropic office at least 25% of the time, with some roles requiring more.
- Competitive benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and office collaboration space.
Tech Stack
Categories
About Anthropic
Anthropic builds large language models and the Claude AI assistant for developers and enterprises, offered via API access and enterprise plans. Founded in 2021 and headquartered in San Francisco, it distributes Claude through its own platform and via partners such as Amazon Bedrock and Google Cloud’s Vertex AI. Its work emphasizes model reliability, interpretability, and practical tooling for tasks like coding assistance, analysis, and customer support automation.
