13 hours ago
Base Salary
$180k - $350k/yr
Responsibilities
- Improve parsing for web pages that current systems cannot process and demonstrate measurable gains.
- Build models that classify page types, assess page quality and credibility, detect misinformation, and identify machine-generated or search-engine-optimized content.
- Design large-scale datasets and supervision for problems without established labels or ground truth.
- Train accurate and cost-efficient classification and extraction models for web-scale deployment.
- Determine whether documents are semantically identical or meaningfully different for web deduplication.
- Trace poor search results to the page-level predictions that caused them and improve the underlying systems.
Requirements
- Graduate-level machine learning experience through a Master’s or PhD with at least 2 years of relevant experience, or an exceptionally strong undergraduate background.
- Ability to build a transformer from scratch in PyTorch.
- Experience training models that are inexpensive enough to run broadly.
- Comfort working with large-scale datasets and defining ground truth for open-ended problems.
- Interest in high-quality knowledge and information retrieval.
Tech Stack
Categories
About Exa
Exa builds a web-scale search engine and API for AI agents and developer applications, combining its own crawler, embedding models, and high-performance vector search. It sells usage-based APIs and enterprise integrations that power retrieval-augmented generation, browsing, and automation on live web content. The company runs large crawling infrastructure and dedicated GPU clusters to continuously index and embed pages, and ships low-latency search components written in Rust.
