2 months ago
Base Salary
$170k - $205k/yr
Responsibilities
- Develop deep-level diagnostics and troubleshooting for hardware faults in GPU racks and high-density compute systems.
- Build troubleshooting and automation tooling for NVIDIA A100, H200, GB200, B200, and AMD 350X/355X GPU platforms.
- Develop automation and AI agents for component-level diagnosis and remediation of failed or degraded hardware.
- Create tooling and AI agents with data center operations teams to manage critical environments.
- Develop post-repair validation and testing tools, including burn-in, PyTorch, and NVIDIA NCCL workflows.
- Own deployment, monitoring, and operational support of developed tooling to maximize GPU fleet availability and performance.
- Develop automation and operational tooling for facilities power and direct liquid cooling hardware systems.
Requirements
- 4–6 years of software engineering experience.
- Expertise in distributed systems, reliability, and cloud platforms such as Kubernetes and GCP.
- Strong programming ability in at least one of Go, Python, Java, or Rust.
- Ability to identify problems, rapidly develop scalable solutions, ship them, and set technical direction for projects.
- Strong analytical, problem-solving, communication, collaboration, and independent-work skills.
- Experience with Temporal and Kubernetes is preferred.
- Experience working directly with hardware vendors or in large-scale GPU fleet operations or hyperscale data center environments is preferred.
Benefits
- Industry-competitive pay and restricted stock units.
- Health, dental, and vision insurance options, including HDHP and PPO plans, plus employer HSA contributions.
- Paid parental leave, paid life insurance, and short- and long-term disability coverage.
- 401(k) with a 100% employer match up to 4% of salary.
- Generous paid time off and holiday schedule.
- Cell phone reimbursement, tuition reimbursement, Teladoc, Calm app subscription, MetLife Legal, and a company-paid commuter benefit of $50 per pay period.
About Crusoe
Crusoe builds and operates GPU-powered data centers and an AI cloud platform for enterprises running large-scale AI training and inference. It vertically integrates energy supply with compute, using stranded and renewable power to reduce emissions and costs, and sells capacity via cloud services and managed infrastructure. Founded in 2018 and headquartered in Denver, the company is privately held.
