5 months ago
Base Salary
$180k - $440k/yr
Responsibilities
- Design, build, and optimize massive GPU clusters for extreme-scale training and inference
- Develop and tune low-level CUDA kernels using GeMM, Attention, CUTLASS, Tensor Cores, and Nsight
- Work on Linux kernel scheduling, memory management, and resource isolation at cluster scale
- Build custom container orchestration, virtualization layers using KVM and Firecracker, and distributed systems beyond standard Kubernetes
- Profile, debug, and eliminate bottlenecks across GPU memory, networking fabric, filesystems, and multi-GPU operations
- Create and maintain infrastructure-as-code, automation, and tools for reliable and efficient supercomputer operations
- Collaborate with AI research teams to deliver production-grade performance and scalability
Requirements
- Deep low-level systems programming experience with C, C++, or Rust
- Experience building and operating high-performance exabyte-scale storage systems
- Strong experience with large-scale GPU clusters or distributed compute infrastructure in production
- Hands-on experience optimizing GPU kernels with CUTLASS, custom kernels, and Nsight profiling
- Experience with Linux kernel internals, scheduling, virtualization, or large-scale orchestration
- A track record of building or operating high-performance infrastructure for AI training or inference workloads
- Ability to reason from first principles and optimize memory-bound and compute-bound workloads
Benefits
- Equity and comprehensive medical, vision, and dental coverage
- Access to a 401(k) retirement plan
- Short- and long-term disability insurance and life insurance
- Various discounts and perks
Tech Stack
Categories
About xAI
Understand the Universe. We are a team of AI technologists and business leaders on a mission to build AI systems that can help humanity understand the world better. https://x.ai/careers