7 hours ago
Base Salary
$180k - $440k/yr
Responsibilities
- Architect and implement distributed model-serving infrastructure for load balancing, auto-scaling, batch scheduling, and global KV caching.
- Optimize model inference latency, throughput, reliability, and tail latency under production workloads.
- Benchmark, fine-tune, and accelerate inference engines, including GPU kernel and code-generation work.
- Develop tools to trace, replay, diagnose, and fix issues across orchestration and GPU-kernel layers.
- Create CI/CD infrastructure for endpoint deployment, image publishing, and inference-engine updates.
- Support research on test-time compute scaling, RL rollout, and model-hardware co-design.
Requirements
- Deep low-level systems programming experience with C, C++, or Rust.
- Experience with large-scale, high-concurrency production serving.
- Experience with GPU inference engines such as vLLM, SGLang, Triton, or TensorRT-LLM.
- Strong background in batching, caching, load balancing, and parallelism optimization.
- Experience with GPU kernels and code generation for inference optimization.
- Experience with quantization, speculative decoding, distillation, and low-precision numerics.
- Experience testing, benchmarking, and ensuring the reliability of inference services.
- Experience designing and implementing CI/CD infrastructure for inference.
Benefits
- Equity and comprehensive medical, vision, and dental coverage.
- Access to a 401(k) retirement plan, short- and long-term disability insurance, life insurance, discounts, and other perks.
Categories
About xAI
Understand the Universe. We are a team of AI technologists and business leaders on a mission to build AI systems that can help humanity understand the world better. https://x.ai/careers