11 days ago
Responsibilities
- Build and own high-throughput, observable data pipelines with end-to-end traceability for frontier-scale data.
- Train and ship models that classify, rank, filter, clean, and identify training data at extreme throughput.
- Design scaling-ladder experiments to evaluate data mixtures, repeatability, and quality.
- Build platforms, orchestration, and tooling that transform raw web, code, multimodal, and acquired data into training-ready datasets.
- Create signals for data quality, lineage, freshness, and pipeline health.
- Build and scale distributed web-crawling systems for discovery, scheduling, fetching, parsing, and ingestion of high-quality documents.
- Improve URL seeding, scoring, host scheduling, crawl success, anti-bot handling, parsing quality, availability, and recovery.
- Partner with data acquisition, data quality, crawling, data platform, and training teams to improve model loss, evaluations, and capability.
Requirements
- Strong infrastructure or data platform background, with crawling or search infrastructure experience considered a plus.
- Ability to architect and ship end-to-end systems with high ownership and independently debug complex systems.
- Strong intuitions about large-scale distributed systems.
- Ability to work alongside AI agents and collaborate across research, training, acquisition, and engineering teams.
- Interest in how pretraining data shapes model quality and in building systems that feed frontier training runs.
- Deep domain experience or unusually rapid engineering growth, such as achieving staff-level ownership within a few years, is considered a fit.
Categories
Data EngineeringML Engineering
