Anthropic

Red Team Engineer, Safeguards

Anthropic
Apply
2 months ago
Remote, Worldwide or San Francisco, CA, USASenior
H1B Sponsor

Base Salary

$320k - $405k/yr

Responsibilities

  • Conduct adversarial testing across product surfaces using creative multi-technique attack scenarios.
  • Research and implement testing approaches for agent systems, tool use, emerging capabilities, and new interaction paradigms.
  • Design and execute full kill-chain attacks that emulate real-world threat actors.
  • Build and maintain systematic testing methodologies for comprehensive system evaluation.
  • Develop automated testing frameworks for continuous assessment at scale.
  • Collaborate with Product, Engineering, and Policy teams to turn findings into concrete improvements.
  • Establish metrics for measuring detection effectiveness against novel abuse.

Requirements

  • Experience in penetration testing, red teaming, or application security.
  • Experience with model jailbreaking and testing large-scale agentic workflows for non-obvious prompt injection vectors.
  • Hands-on expertise in web application security and security testing tools such as Burp Suite and Metasploit.
  • Experience building custom automation, including LLM-specific testing frameworks.
  • Track record of discovering novel attack vectors and creatively chaining vulnerabilities.
  • Public body of work such as CVEs, blog posts, or disclosed bug bounty reports.
  • Strong written and verbal communication skills.
  • Preferred experience in AI/ML security or adversarial machine learning.
  • Preferred understanding of AI safety considerations and guardrails against jailbreaks.
  • Preferred experience testing API security and rate-limiting systems.
  • Preferred background in business logic vulnerabilities, authorization bypass techniques, anti-fraud, trust and safety, or abuse prevention.
  • Preferred familiarity with distributed systems, infrastructure security, and abuse detection mechanisms.

Benefits

  • Hybrid work policy requiring staff to be in an office at least 25% of the time
  • Visa sponsorship may be available
  • Competitive compensation and benefits
  • Optional equity donation matching
  • Generous vacation and parental leave
  • Flexible working hours
  • Office space for collaboration

Categories

Anthropic

About Anthropic

501-1,000 employees

We're an AI research company that builds reliable, interpretable, and steerable AI systems. Our first product is Claude, an AI assistant for tasks at any scale. Our research interests span multiple areas including natural language, human feedback, scaling laws, reinforcement learning, code generation, and interpretability.