
Member of Technical Staff - Research Software Engineer - Safety Evaluations Infrastructure
Reflection2 months ago
London, United Kingdom +2 moreSenior
Responsibilities
- Design and build secure, sandboxed infrastructure for sensitive model evaluations, including CBRN and other dangerous-capability domains.
- Build controlled data pipelines and storage with least-privilege access, RBAC, encryption, audit logging, and data-minimization safeguards.
- Translate evaluation designs into reliable, reproducible, and scalable systems with safety researchers and domain experts.
- Develop evaluation orchestration tooling and harnesses for high-throughput model and agent evaluations in isolated environments.
- Build infrastructure to measure AI capability uplift and integrate results into release-decision pipelines.
- Implement guardrails, monitoring, and compartmentalization across compute, data, and tooling.
- Write production-quality Python and related tooling for high-throughput data processing and evaluation systems.
- Improve the reliability, security posture, and developer experience of the safety evaluation platform.
Requirements
- Strong software engineering skills, particularly in Python, with experience building reliable, scalable infrastructure or platform systems.
- Experience building sandboxed, isolated, or security-sensitive execution environments using approaches such as containerization, VM isolation, or secure compute.
- Grounding in security engineering fundamentals including least privilege, need-to-know access, RBAC, secrets management, encryption, audit logging, compartmentalization, and defense in depth.
- Experience building data pipelines and handling sensitive or restricted data with appropriate safeguards.
- Ability to own ambiguous, cross-functional problems end to end.
- Discretion, integrity, sound judgment, and comfort working on sensitive projects.
- Experience with evaluation, benchmarking, or experimentation infrastructure for ML systems is advantageous.
- Experience with LLMs, agents, or ML training and inference pipelines is advantageous.
- Familiarity with dangerous-capability or dual-use domains and information-security considerations is advantageous.
- Familiarity with compliance frameworks relevant to sensitive data handling is advantageous.
Benefits
- Top-tier compensation and equity structure, including stock options.
- Comprehensive medical, dental, vision, and life insurance with an annual wellness allowance.
- Lunch and dinner provided daily in the office.
- 22 weeks of paid parental leave for birthing and non-birthing parents, including adoptive and surrogate journeys.
- Unlimited paid time off in the U.S. and 30 vacation days in the U.K.
- Visa sponsorship and support for long-term immigration pathways where applicable.
- Regular off-sites, happy hours, and team celebrations.
About Reflection
Reflection is a New York–based, privately held research lab developing open foundational AI models and agentic coding tools for developers, enterprises, and public-sector users. The team includes former researchers from DeepMind, OpenAI, and Anthropic, and their work focuses on transparent, customizable systems that organizations can deploy with ownership and control.