22 days ago
Base Salary
$215k - $260k/yr
Responsibilities
- Develop deep-level diagnostics and troubleshooting for hardware faults in GPU racks and high-density compute systems.
- Build troubleshooting and automation tooling for NVIDIA and AMD GPU platforms.
- Develop automation and AI agents for component-level diagnosis and remediation of failed or degraded hardware.
- Create tooling with data center operations teams for critical-environment management.
- Develop post-repair validation and testing tools, including burn-in, PyTorch, and NVIDIA NCCL workflows.
- Own deployment, monitoring, and operational support of developed tooling to improve GPU fleet availability and performance.
- Develop automation and operational tooling for facility power management and direct liquid-cooling hardware systems.
- Set technical direction for projects and execute scalable solutions independently and collaboratively.
Requirements
- Professional software engineering experience and the ability to rapidly develop and ship scalable solutions.
- Expertise in distributed systems, reliability, and cloud platforms such as Kubernetes and GCP.
- Strength in at least one of Go, Python, Java, or Rust.
- Strong analytical, problem-solving, communication, collaboration, and independent execution skills.
- Preferred: experience with Temporal and Kubernetes.
- Preferred: experience working directly with hardware vendors.
- Preferred: background in large-scale GPU fleet operations or hyperscale data center environments.
Benefits
- Industry-competitive pay, bonus eligibility, and restricted stock units.
- Health insurance options including HDHP and PPO, plus vision and dental coverage for employees and dependents.
- Employer HSA contributions, paid parental leave, paid life insurance, and short- and long-term disability coverage.
- Teladoc, a 401(k) with a 100% match up to 4% of salary, paid time off, and a holiday schedule.
- Cell phone reimbursement, tuition reimbursement, Calm app subscription, MetLife Legal, and a company-paid commuter benefit of $300 per month.
About Crusoe
As the AI factory company, Crusoe is on a mission to accelerate the abundance of energy and intelligence. The company provides a reliable, scalable, cost-effective, and energy-first solution for AI infrastructure. By harnessing large-scale energy resources, building AI-optimized data centers, and delivering an AI cloud platform, Crusoe empowers its customers to build the future faster.
