12 months ago
Base Salary
$180k - $350k/yr
Responsibilities
- Build Kubernetes orchestration for a $20 million GPU cluster.
- Scale the AWS Batch system to support map-reduce jobs across tens of thousands of machines.
- Design GPU scheduling software to maximize cluster utilization.
- Build observability into production systems.
Requirements
- Experience designing and operating large-scale infrastructure, such as GPU clusters, large Kubernetes clusters, or cloud batch-job systems.
- Strong focus on reliability, observability, and optimization across the infrastructure stack.
Benefits
- In-person role in San Francisco.
- International visa sponsorship is available, including STEM OPT, OPT, H1B, O1, and E3.
Tech Stack
Categories
About Exa
Exa is an applied AI research lab organizing human knowledge, starting with the web. Exa Search is the highest quality search API at every latency. It accepts natural language queries and produces token-efficient results with citations. Exa Agent is a high-compute web agent that synthesizes information from across the web to handle the most complex deep research, list-building, and enrichment workflows. These products power Cursor, Cognition, Hubspot, OpenRouter, Monday.com, and agents built by over 400,000 developers. We are a worldwide team of engineers and researchers. We're hiring: exa.ai/careers
