1 month ago
Remote, United StatesSenior / Staff+
Responsibilities
- Build and optimize LLM serving and inference systems for production environments.
- Improve performance across GPU and CPU pathways and reduce throughput and latency bottlenecks.
- Work on KV cache, memory, storage, and serving architecture challenges.
- Design and scale systems supporting RAG and retrieval-heavy AI workloads.
- Contribute to infrastructure where storage architecture and systems efficiency affect AI performance.
- Solve problems at the intersection of AI, high-performance systems, and distributed infrastructure.
- Move between system architecture decisions and hands-on implementation.
Requirements
- Meaningful experience building or optimizing production AI systems rather than only experimenting with models.
- Deep hands-on systems-layer experience involving GPU and CPU resources, infrastructure tuning, bottleneck reduction, throughput, or latency.
- Evidence of ownership in model serving, retrieval, caching, storage, or distributed performance.
- Ability to work across architecture and implementation in environments where efficiency and scale matter.
- Background in AI infrastructure, high-performance systems, storage platforms, distributed systems, or an adjacent area.
- Purely academic research experience without meaningful production ownership is not sufficient.
- PhD preferred, but real-world systems experience is valued more highly.