6 months ago
Base Salary
$320k - $405k/yr
Responsibilities
- Build systems that gather and analyze signals at scale to identify and prevent account abuse.
- Integrate third-party data-enrichment vendors.
- Inspect other teams’ code to identify signal-collection points and introduce low-impact interventions.
- Create monitoring dashboards, alerts, and internal administrative UX.
- Collaborate with data scientists on usage patterns and trends and with Policy & Enforcement on human-review effectiveness.
- Build robust, reliable, multilayered defenses.
- Lead root cause analyses and investigations into account activity, abuse patterns, and emerging attack vectors.
- Inform immediate enforcement actions and longer-term systemic defenses.
Requirements
- Bachelor’s degree in Computer Science, Software Engineering, or a comparable field, or equivalent experience.
- 3–10+ years of experience in a software engineering position, preferably focused on integrity, spam, fraud, or abuse detection.
- Proficiency in Python, SQL, and data analysis tools.
- Strong communication skills and the ability to explain complex technical concepts to non-technical stakeholders.
- Experience with trust and safety mechanisms, AI/ML systems, fraud-detection models, security monitoring tools, or supporting infrastructure is a strong preferred qualification.
- Experience working with operational teams to build custom internal tooling is a strong preferred qualification.
Benefits
- Annual salary range of $320,000–$405,000 USD.
- Hybrid policy requiring staff to be in an office at least 25% of the time.
- Visa sponsorship may be available, with immigration lawyer support.
- Competitive benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and office collaboration space.
About Anthropic
We're an AI research company that builds reliable, interpretable, and steerable AI systems. Our first product is Claude, an AI assistant for tasks at any scale. Our research interests span multiple areas including natural language, human feedback, scaling laws, reinforcement learning, code generation, and interpretability.