5 hours ago
Base Salary
$180k - $350k/yr
Responsibilities
- Own the signals and decisions determining which pages are indexed, refreshed, followed, deduplicated, or excluded.
- Develop and apply classifiers, rankers, and calibration methods for web-scale indexing decisions.
- Define what quality, credibility, semantic similarity, and indexing value mean in problems without clear ground truth.
- Measure end-to-end impact by proving that modeling improvements change decisions and improve search.
- Collaborate with crawling, indexing, and retrieval teams that consume the resulting systems.
- Improve page parsing coverage and model page quality, misinformation, source reliability, and crawler-oriented content.
- Determine whether documents are semantically equivalent or meaningfully distinct for web deduplication.
Requirements
- Hands-on machine-learning experience with classifiers, rankers, and calibration.
- Strong data instincts and experience working at web scale.
- Ability to own ambiguous problems where defining the correct answer is part of the work.
- End-to-end product and systems thinking focused on measurable search impact.
- Ability and interest in collaborating across crawling, indexing, and retrieval teams.
- Interest in finding high-quality knowledge and improving search over the world’s information.
Categories
About Exa
Exa builds a web-scale search engine and API for AI agents and developer applications, combining its own crawler, embedding models, and high-performance vector search. It sells usage-based APIs and enterprise integrations that power retrieval-augmented generation, browsing, and automation on live web content. The company runs large crawling infrastructure and dedicated GPU clusters to continuously index and embed pages, and ships low-latency search components written in Rust.
