2 months ago
Responsibilities
- Design and optimize the inference stack for RL workloads ranging from small-scale ablations to production training runs.
- Analyze, profile, and resolve performance bottlenecks in large-scale RL systems.
- Collaborate with the modeling team to efficiently implement novel RL techniques and algorithms.
Requirements
- Experience building, debugging, and optimizing the efficiency of large-scale distributed systems.
- Experience with LLM inference.
- Proficiency in Python, C++, or Rust and frameworks such as PyTorch, JAX, or CUDA.
- Strong knowledge of quantization and numerics in LLM inference and training is preferred.
- Experience developing inference engines such as SGLang or vLLM is preferred.
- Willingness to solve complex problems across all levels of the stack.
About xAI
Understand the Universe. We are a team of AI technologists and business leaders on a mission to build AI systems that can help humanity understand the world better. https://x.ai/careers