4 hours ago
Base Salary
$250k - $300k/yr
Responsibilities
- Develop deep-level diagnostics and troubleshooting for hardware faults in GPU racks and high-density compute systems.
- Build troubleshooting and automation tooling for NVIDIA A100, H200, GB200, B200, and AMD 350X and 355X GPU platforms.
- Develop automation and AI agents for component-level diagnosis and remediation of failed or degraded hardware.
- Create tooling and AI agents with data center operations teams for managing critical environments.
- Develop post-repair validation and testing tools, including burn-in, PyTorch, and NVIDIA NCCL, to ensure system stability and performance.
- Own deployment, monitoring, and operational support of tooling to maximize GPU fleet availability and performance.
- Develop automation and operational tooling for facilities power management and direct liquid cooling hardware systems.
Requirements
- Software engineering experience and the ability to rapidly develop and ship scalable solutions.
- Expertise in distributed systems, reliability, and cloud platforms including Kubernetes and GCP.
- Strong proficiency in at least one of Go, Python, Java, or Rust.
- Ability to set technical direction for projects, work independently, and collaborate on critical technical initiatives.
- Strong analytical, problem-solving, communication, and collaboration skills.
- Experience with Temporal and Kubernetes is preferred.
- Experience working directly with hardware vendors is preferred.
- Background in large-scale GPU fleet operations or hyperscale data center environments is preferred.
Benefits
- Industry-competitive pay, restricted stock units, health insurance with HDHP and PPO options, vision and dental coverage, and employer HSA contributions.
- Paid parental leave, paid life insurance, short- and long-term disability coverage, Teladoc, and a 401(k) with a 100% match up to 4% of salary.
- Generous paid time off and holidays, cell phone reimbursement, tuition reimbursement, Calm app subscription, MetLife Legal, and a company-paid commuter benefit of $300 per month.
About Crusoe
As the AI factory company, Crusoe is on a mission to accelerate the abundance of energy and intelligence. The company provides a reliable, scalable, cost-effective, and energy-first solution for AI infrastructure. By harnessing large-scale energy resources, building AI-optimized data centers, and delivering an AI cloud platform, Crusoe empowers its customers to build the future faster.
