Cursor

Software Engineer, Pretraining

Cursor
Apply
11 days ago

Responsibilities

  • Build and own high-throughput, observable data pipelines with end-to-end traceability for frontier-scale data.
  • Train and ship models that classify, rank, filter, clean, and identify training data at extreme throughput.
  • Design scaling-ladder experiments to evaluate data mixtures, repeatability, and quality.
  • Build platforms, orchestration, and tooling that transform raw web, code, multimodal, and acquired data into training-ready datasets.
  • Create signals for data quality, lineage, freshness, and pipeline health.
  • Build and scale distributed web-crawling systems for discovery, scheduling, fetching, parsing, and ingestion of high-quality documents.
  • Improve URL seeding, scoring, host scheduling, crawl success, anti-bot handling, parsing quality, availability, and recovery.
  • Partner with data acquisition, data quality, crawling, data platform, and training teams to improve model loss, evaluations, and capability.

Requirements

  • Strong infrastructure or data platform background, with crawling or search infrastructure experience considered a plus.
  • Ability to architect and ship end-to-end systems with high ownership and independently debug complex systems.
  • Strong intuitions about large-scale distributed systems.
  • Ability to work alongside AI agents and collaborate across research, training, acquisition, and engineering teams.
  • Interest in how pretraining data shapes model quality and in building systems that feed frontier training runs.
  • Deep domain experience or unusually rapid engineering growth, such as achieving staff-level ownership within a few years, is considered a fit.

Categories

Data EngineeringML Engineering
Cursor

About Cursor

1,001-5,000 employees
Contact me