8 months ago
Base Salary
$180k - $220k/yr
Responsibilities
- Serve as the highest-level escalation point for complex P1/P0 incidents
- Lead cross-functional root-cause investigations across compute, networking, storage, and orchestration layers
- Partner with SRE and Storage, Networking, Compute, and Kubernetes software teams on systemic fixes
- Design and improve node validation, burn-in, performance baselining, and release-readiness processes
- Influence Kubernetes architecture, Slurm and Terraform workload orchestration, and AI/ML cluster stability
- Reduce MTTR and incident recurrence through structural reliability improvements
- Troubleshoot NCCL, InfiniBand, GPU driver and firmware issues, and distributed training failures
- Tune performance and improve observability for large-scale AI training and inference workloads
- Advise customers during high-risk incidents and deliver executive-ready root-cause analyses
- Mentor P3/P4 engineers and define SOPs and technical standards for support excellence
Requirements
- 8+ years of experience in SRE, DevOps, HPC, or cloud infrastructure roles
- Advanced Linux systems expertise
- Deep Kubernetes operational experience at CKA level or higher
- Strong networking knowledge including InfiniBand, RDMA, RoCE, and SDN
- Experience supporting AI/ML workloads at scale on GPU clusters
- Proven ability to resolve multi-layer distributed system failures
- Strong customer communication and executive-facing presence
Benefits
- Restricted Stock Units
- Paid time off and paid holidays
- Comprehensive health, dental, and vision insurance
- Employer HSA contributions
- Paid parental leave
- Paid life insurance and short-term and long-term disability coverage
- Professional development and tuition reimbursement
- Mental health and wellness support
- Commuter benefits for parking and transit
- Cell phone stipend
- 401(k) retirement plan with company match up to 4% of salary
- Volunteer time off
Tech Stack
Categories
DevOpsSite Reliability
About Crusoe
Crusoe builds and operates GPU-powered data centers and an AI cloud platform for enterprises running large-scale AI training and inference. It vertically integrates energy supply with compute, using stranded and renewable power to reduce emissions and costs, and sells capacity via cloud services and managed infrastructure. Founded in 2018 and headquartered in Denver, the company is privately held.
