2 months ago
Palo Alto, CA, USAMid Level
H1B sponsor
Base Salary
$180k - $440k/yr
Responsibilities
- Design, build, and scale GPU compute infrastructure, training frameworks, and experimentation tools for rapid ML iteration.
- Develop data pipelines and integrate large-scale data, training, and inference systems.
- Collaborate with ML teams to productionize models and ensure seamless integration across the stack.
- Improve the scalability, reliability, and efficiency of large-scale machine-learning systems.
- Work across the full stack to solve complex problems independently.
- Mentor junior engineers and contribute to team growth.
Requirements
- Bachelor's, master's, postgraduate, or PhD in computer science, machine learning, or another quantitative discipline, or equivalent work experience.
- At least two years of industry experience with high-traffic or large-scale production environments, distributed systems, GPU infrastructure, and/or deep learning applications.
- At least two years of experience with ML platforms, training infrastructure, or close collaboration with modeling engineers and data scientists.
- Strong proficiency with Python and experience with compiled languages such as C++ or Rust.
- Deep familiarity with modern ML frameworks such as JAX or PyTorch is preferred.
- Low-level understanding of distributed storage, NVIDIA drivers, CUDA toolkits, and networking is preferred.
- Experience with Linux systems, orchestration tools, job schedulers such as Slurm, configuration management tools such as Puppet or Ansible, or related infrastructure tooling is preferred.
Benefits
- Equity and comprehensive medical, vision, and dental coverage.
- Access to a 401(k) retirement plan, short- and long-term disability insurance, life insurance, discounts, and other perks.
About xAI
Understand the Universe. We are a team of AI technologists and business leaders on a mission to build AI systems that can help humanity understand the world better. https://x.ai/careers