Exa

Software Engineer, Distributed Data Systems

Exa
Apply
10 months ago
San Francisco, CA, USAMid Level / Senior
H1B sponsor

Base Salary

$180k - $350k/yr

Responsibilities

  • Design a lakehouse architecture that handles over 100 PB of web crawl data.
  • Build streaming pipelines that process billions of documents daily for real-time indexing.
  • Architect the data layer for embedding training infrastructure on Ray.
  • Scale ClickHouse deployment to manage analytical queries across petabytes of search logs.

Requirements

  • Deep understanding of lakehouse architectures like Delta Lake, Iceberg, and Hudi.
  • Experience in building and operating large-scale distributed data processing pipelines.
  • Hands-on experience with streaming data systems such as Kafka or Flink.
  • Familiarity with Ray, Spark, or ClickHouse at production scale.
  • Focus on reliability to build systems that minimize operational issues.

Benefits

  • Premium healthcare benefits including medical, dental, and vision.
  • Fertility benefits offered to all employees.
  • Monthly wellness stipend provided.

Tech Stack

Apache FlinkApache KafkaApache SparkClickHouseRust

Categories

Data Engineering
Exa

About Exa

51-200 employees

Exa builds a web-scale search engine and API for AI agents and developer applications, combining its own crawler, embedding models, and high-performance vector search. It sells usage-based APIs and enterprise integrations that power retrieval-augmented generation, browsing, and automation on live web content. The company runs large crawling infrastructure and dedicated GPU clusters to continuously index and embed pages, and ships low-latency search components written in Rust.

Contact me