6 months ago
Base Salary
$180k - $220k/yr
Responsibilities
- Serve as the highest-level escalation point for complex P1/P0 incidents
- Lead cross-functional root-cause investigations across compute, networking, storage, and orchestration layers
- Partner with SRE and Storage, Networking, Compute, and Kubernetes software teams on systemic fixes
- Design and improve node validation, burn-in, performance baselining, and release-readiness processes
- Influence Kubernetes architecture, Slurm and Terraform workload orchestration, and AI/ML cluster stability
- Reduce MTTR and incident recurrence through structural reliability improvements
- Troubleshoot NCCL, InfiniBand, GPU driver and firmware issues, and distributed training failures
- Tune performance and improve observability for large-scale AI training and inference workloads
- Advise customers during high-risk incidents and deliver executive-ready root-cause analyses
- Mentor P3/P4 engineers and define SOPs and technical standards for support excellence
Requirements
- 8+ years of experience in SRE, DevOps, HPC, or cloud infrastructure roles
- Advanced Linux systems expertise
- Deep Kubernetes operational experience at CKA level or higher
- Strong networking knowledge including InfiniBand, RDMA, RoCE, and SDN
- Experience supporting AI/ML workloads at scale on GPU clusters
- Proven ability to resolve multi-layer distributed system failures
- Strong customer communication and executive-facing presence
Benefits
- Restricted Stock Units
- Paid time off and paid holidays
- Comprehensive health, dental, and vision insurance
- Employer HSA contributions
- Paid parental leave
- Paid life insurance and short-term and long-term disability coverage
- Professional development and tuition reimbursement
- Mental health and wellness support
- Commuter benefits for parking and transit
- Cell phone stipend
- 401(k) retirement plan with company match up to 4% of salary
- Volunteer time off
Tech Stack
Categories
DevOpsSite Reliability
About Crusoe
As the AI factory company, Crusoe is on a mission to accelerate the abundance of energy and intelligence. The company provides a reliable, scalable, cost-effective, and energy-first solution for AI infrastructure. By harnessing large-scale energy resources, building AI-optimized data centers, and delivering an AI cloud platform, Crusoe empowers its customers to build the future faster.
