Reflection

Member of Technical Staff - Safety Lead

Reflection
Apply
9 months ago
London, United Kingdom +2 moreStaff+
H1B sponsor

Responsibilities

  • Own the red-teaming and adversarial evaluation pipeline for Reflection’s models across security, misuse, and alignment failure modes.
  • Collaborate with the Alignment team to translate safety findings into concrete guardrails and deployment policies.
  • Validate that model releases meet the lab’s risk thresholds before shipping.
  • Develop scalable automated safety benchmarks using dynamic adversarial testing.
  • Research and implement state-of-the-art jailbreaking techniques and defenses.

Requirements

  • Graduate degree (MS or PhD) in Computer Science, Machine Learning, or a related discipline, or equivalent practical experience in AI Safety.
  • Deep technical understanding of LLM safety, adversarial attacks, red-teaming methodologies, and interpretability.
  • Strong software engineering capabilities and experience building automated evaluation pipelines or large-scale ML systems.
  • Experience with Reinforcement Learning from Human Feedback (RLHF) and Reinforcement Learning from AI Feedback (RLAIF) is a strong plus.
  • Ability to make high-stakes decisions regarding model releases and safety thresholds.

Benefits

  • Top-tier compensation with salary and equity structured to recognize and retain talent globally.
  • Comprehensive medical, dental, vision, life, and disability insurance.
  • Fully paid parental leave for all new parents, including adoptive and surrogate journeys.
  • Financial support for family planning.
  • Paid time off and relocation support.
  • Daily lunch and dinner, regular off-sites, and team celebrations.

Categories

AI ResearchTesting
Reflection

About Reflection

201-500 employees

Reflection is a New York–based, privately held research lab developing open foundational AI models and agentic coding tools for developers, enterprises, and public-sector users. The team includes former researchers from DeepMind, OpenAI, and Anthropic, and their work focuses on transparent, customizable systems that organizations can deploy with ownership and control.

Contact me