2 months ago
Base Salary
$180k - $440k/yr
Responsibilities
- Design, build, and implement large-scale distributed systems powering a supercomputing cluster.
- Profile, debug, and optimize performance across GPUs, the Linux kernel, networking, and filesystems.
- Collaborate on hardware, software, and algorithm co-design for AI training.
- Maintain and improve the codebase for scalability and reliability.
- Develop tools that improve team productivity and streamline workflows.
Requirements
- Systems programming experience in C, C++, or Rust.
- Strong computer systems fundamentals, from transistors through high-level applications.
- Hands-on Kubernetes expertise, including cluster architecture, pod lifecycle, CNI networking, CSI storage, service mesh, and production-grade operations.
- Preferred experience with operating systems internals, networking and TCP/IP, performance analysis, low-level optimization, and full-stack debugging.
- Preferred experience with Linux systems-level debugging tools such as perf, gdb, strace, or Wireshark.
- Preferred experience deploying and managing Kubernetes workloads with manifests, Helm, Operators, and GitOps workflows.
- Preferred understanding of Docker, containerd, and CRI-O and their interaction with the Linux kernel.
- Preferred experience with observability and monitoring tools such as Prometheus, Grafana, VictoriaMetrics, or OpenTelemetry.
Benefits
- Equity compensation
- Comprehensive medical, vision, and dental coverage
- 401(k) retirement plan
- Short- and long-term disability insurance
- Life insurance
- Various discounts and perks
About xAI
Understand the Universe. We are a team of AI technologists and business leaders on a mission to build AI systems that can help humanity understand the world better. https://x.ai/careers