7 months ago
Remote, United StatesSenior
Responsibilities
- Track current research, open-source datasets, and models related to LLMs and synthetic data generation.
- Design and implement complex, cost-efficient pipelines that generate diverse, high-quality datasets at scale.
- Collaborate cross-functionally to ensure experiments and generated data use compute and time efficiently.
- Measure and refine dataset quality and validate data strategies through quantitative ablation experiments.
- Lead original, time-bounded research initiatives and deploy technical engineering solutions into production.
Requirements
- Strong machine learning and engineering background with experience working with LLMs and how they learn.
- Knowledge of data ablations, scaling laws, post-training techniques, and training reasoning and agentic models.
- Experience generating synthetic datasets at scale while optimizing quality, correctness, diversity, and cost.
- Experience with model capability evaluations covering areas such as knowledge, reasoning, mathematics, coding, and long context.
- Experience building trillion-scale pre-training datasets and knowledge of curation, deduplication, data mixing, tokenization, curriculum, and data repetition.
- Excellent Python programming skills and strong prompt-engineering skills.
- Experience with large-scale GPU clusters and distributed data pipelines.
- Research experience and the ability to discuss recent papers in technical detail.
- Preferred qualifications include scientific publications in applied deep learning, LLMs, or source-code generation, and formal machine learning, mathematics, or computer science training.
Benefits
- Fully remote work with flexible hours.
- 37 days per year of vacation and holidays.
- Health insurance allowance for the employee and dependents.
- 16 weeks of flexible, fully paid parental leave.
- Company-provided equipment.
- Well-being, continuous-learning, and home-office allowances.
- Frequent team gatherings and a diverse, inclusive, people-first culture.
- The distributed team meets in Paris monthly for three days, with a lower cadence discussable for people based in PST, plus annual off-sites.
Tech Stack
Categories
AI ResearchML Engineering
