
Sieve
Sieve builds the data and environments frontier AI labs use to train the next generation of multimodal systems. AI is moving beyond chatbots into video, audio, images, software, robotics, and interactive worlds. The next generation of models will need to understand how the world looks, sounds, moves, responds, and changes over time. Progress is bottlenecked by one thing: high-quality data. Sieve brings together exabyte-scale infrastructure, novel multimodal understanding techniques, large-scale sourcing, and deep research partnerships to create datasets and environments with unmatched precision, quality, and speed. This has earned the trust of frontier AI labs, Fortune 100 companies, and fast-growing AI startups working on generative media, robotics, computer use, world models, and agentic systems.
Open Positions at Sieve
5 open positions
Build end-to-end machine learning systems for high-quality multimodal datasets and customer-facing video understanding. You will fine-tune foundation models, create evaluation and QA pipelines, and ship production improvements directly to frontier AI labs.
Build custom algorithms, models, and production pipelines that help frontier AI labs create and evaluate high-quality video datasets. This customer-facing, forward-deployed role combines machine learning engineering, research prototyping, and fast delivery of reliable systems.
Build Sieve's full-stack video collection platform, including web, backend, systems, mobile capture, and internal tooling. You'll own projects end-to-end and help power the training data infrastructure behind frontier video AI.
Build and scale the data pipelines, ML filters, and internal tools behind high-quality video datasets for leading AI labs. This end-to-end software engineering role combines broad technical ownership with direct customer collaboration in Sieve’s San Francisco office.
Build high-performance distributed systems that orchestrate ML and ETL pipelines across petabytes of video data. This individual contributor role emphasizes cloud infrastructure, reliability, large-scale GPU systems, and Go and Python development.