
Member of Technical Staff - Safety Lead
Reflection9 months ago
Responsibilities
- Own the red-teaming and adversarial evaluation pipeline for Reflection’s models across security, misuse, and alignment failure modes.
- Collaborate with the Alignment team to translate safety findings into concrete guardrails and deployment policies.
- Validate that model releases meet the lab’s risk thresholds before shipping.
- Develop scalable automated safety benchmarks using dynamic adversarial testing.
- Research and implement state-of-the-art jailbreaking techniques and defenses.
Requirements
- Graduate degree (MS or PhD) in Computer Science, Machine Learning, or a related discipline, or equivalent practical experience in AI Safety.
- Deep technical understanding of LLM safety, adversarial attacks, red-teaming methodologies, and interpretability.
- Strong software engineering capabilities and experience building automated evaluation pipelines or large-scale ML systems.
- Experience with Reinforcement Learning from Human Feedback (RLHF) and Reinforcement Learning from AI Feedback (RLAIF) is a strong plus.
- Ability to make high-stakes decisions regarding model releases and safety thresholds.
Benefits
- Top-tier compensation with salary and equity structured to recognize and retain talent globally.
- Comprehensive medical, dental, vision, life, and disability insurance.
- Fully paid parental leave for all new parents, including adoptive and surrogate journeys.
- Financial support for family planning.
- Paid time off and relocation support.
- Daily lunch and dinner, regular off-sites, and team celebrations.
Categories
AI ResearchTesting
About Reflection
Reflection is a New York–based, privately held research lab developing open foundational AI models and agentic coding tools for developers, enterprises, and public-sector users. The team includes former researchers from DeepMind, OpenAI, and Anthropic, and their work focuses on transparent, customizable systems that organizations can deploy with ownership and control.