11 months ago
San Francisco, CA, USA or New York, NY, USAMid Level
H1B Sponsor
Base Salary
$315k - $425k/yr
Responsibilities
- Build machine learning models to detect unwanted or anomalous behavior from users and API partners and integrate them into production systems.
- Improve automated detection and enforcement systems.
- Analyze reports of inappropriate accounts and build models to proactively detect similar cases.
- Identify abuse patterns and surface them to research teams to improve model training and safety.
Requirements
- At least 4 years of experience in research/ML engineering or an applied research scientist role, preferably focused on AI safety.
- Bachelor’s degree in a related field or equivalent experience.
- Proficiency in Python, LLMs, SQL, and data analysis or data mining tools.
- Experience building safe AI/ML systems such as behavioral classifiers or anomaly detection systems.
- Strong communication skills and the ability to explain complex technical concepts to non-technical stakeholders.
- Experience with Scikit-Learn, TensorFlow, or PyTorch is advantageous.
- Experience with high-performance, large-scale ML systems, transformer language modeling, reinforcement learning, or large-scale ETL is advantageous.
- Interest in the societal impacts and long-term implications of AI work.
Benefits
- Annual base compensation of $315,000-$425,000 USD, plus equity, benefits, and possible incentive compensation.
- Hybrid policy requiring staff to work from an office at least 25% of the time, with some roles requiring more.
- Visa sponsorship and immigration-lawyer support may be available.
- Optional equity donation matching, generous vacation and parental leave, flexible working hours, and office collaboration space.
Categories
About Anthropic
We're an AI research company that builds reliable, interpretable, and steerable AI systems. Our first product is Claude, an AI assistant for tasks at any scale. Our research interests span multiple areas including natural language, human feedback, scaling laws, reinforcement learning, code generation, and interpretability.