
Member of Technical Staff - Multilingual Data
Reflection8 hours ago
London, United Kingdom +2 moreSenior
Responsibilities
- Design and operate large-scale multilingual data pipelines for sourcing, cleaning, deduplication, language identification, and script normalization.
- Define and enforce quality bars for translation quality, cultural fidelity, toxicity, and contamination in multilingual corpora.
- Design and run experiments studying how to improve multilingual data efficiency for large language models.
- Lead small research projects independently while collaborating on larger initiatives.
- Build evaluation sets and diagnostics to identify model degradation by language, register, or domain, and address gaps with targeted data.
- Collaborate with pre-training, mid-training, and post-training teams to improve multilingual model capability.
Requirements
- Strong software engineering fundamentals and experience processing web-scale datasets in distributed environments.
- Experience building large-scale data pipelines for language models, machine translation, speech, or search, ideally across more than one language.
- Fluency or working proficiency in at least one language other than English.
- A rigorous, measurement-first approach to validating data changes.
- Ability to balance research objectives with practical engineering constraints and work with ambiguity and high ownership.
Benefits
- Top-tier compensation and equity, including stock options; compensation figures are not specified.
- Comprehensive medical, dental, vision, and life insurance with an annual wellness allowance.
- Lunch and dinner provided daily in the office.
- 22 weeks of paid parental leave for birthing and non-birthing parents, including adoptive and surrogate journeys.
- Unlimited paid time off in the U.S. and 30 vacation days in the U.K.
- Visa sponsorship and support for long-term immigration pathways where applicable.
- Regular off-sites, happy hours, and team celebrations.
Categories
Data EngineeringML Engineering
About Reflection
Reflection is a New York–based, privately held research lab developing open foundational AI models and agentic coding tools for developers, enterprises, and public-sector users. The team includes former researchers from DeepMind, OpenAI, and Anthropic, and their work focuses on transparent, customizable systems that organizations can deploy with ownership and control.