7 months ago
Remote, EMEA or London, United KingdomSenior
Responsibilities
- Build and maintain high-performance data pipelines processing trillions of tokens.
- Deliver diverse, high-quality datasets for pretraining foundation models and coding agents.
- Engineer ingestion, deduplication, streaming, data modeling, algorithmic sorting, and distributed pipeline optimization at petabyte scale.
- Collaborate with Pretraining, Posttraining, Evals, and Product teams to align dataset quality with model capabilities and downstream use cases.
Requirements
- Strong experience building production-grade distributed data systems for machine learning.
- Experience with orchestration tools such as Slurm, Airflow, or Dagster.
- Experience with observability and reliability tools such as Grafana and Prometheus.
- Experience with Git, Docker, Kubernetes, and cloud-managed services.
- Experience with batch inference, such as vLLM, and large-scale GPU clusters or distributed pipelines.
- Expert-level Python knowledge, strong algorithmic foundations, and the ability to write clean, maintainable code.
- Proficiency with Polars, Dask, or PySpark.
- Preferred experience includes building trillion-scale state-of-the-art pretraining datasets, translating research to production at scale, OCR, web crawling, evaluations, or pretraining large language models.
Benefits
- Fully remote work with flexible hours.
- 37 days per year of vacation and holidays.
- Health insurance allowance for the employee and dependents.
- Company-provided equipment.
- Well-being, continuous-learning, and home-office allowances.
- Frequent team gatherings, including three days of in-person collaboration in Paris each month and annual off-sites.
- Diverse and inclusive people-first culture.
Tech Stack
Categories
Data EngineeringML Engineering
