2 months ago
Base Salary
$170k - $205k/yr
Responsibilities
- Build automation and self-healing tools to monitor and maintain distributed block, file, and object storage infrastructure.
- Drive reliability initiatives covering data replication, encryption, backup and restore, failover, availability, performance, and error budgets.
- Implement and maintain high-performance NVMe- and SSD-backed volumes for large-scale AI compute clusters.
- Investigate and resolve storage incidents using telemetry, logs, and performance profiling.
- Partner with hardware and kernel teams to diagnose low-level I/O issues and optimize I/O paths, cache policies, and file systems.
- Contribute to fault-tolerant, scalable storage backend architecture for AI-focused cloud environments.
Requirements
- Bachelor’s degree in Computer Science, Electrical Engineering, or a related technical field, or equivalent practical experience.
- At least 5 years of professional experience in Storage SRE, systems, or storage engineering.
- Hands-on experience architecting and operating enterprise storage platforms such as Pure Storage or EMC.
- Deep understanding of object, block, and file storage paradigms, Linux internals, I/O subsystems, memory management, and storage scheduling.
- Proficiency in Go, Python, Java, or C.
- Experience with Infrastructure as Code and deployment tooling such as Terraform, Ansible, or Puppet.
- Familiarity with NFS, SMB, iSCSI, or NVMe-oF storage protocols.
- Strong experience with containerized workloads and orchestration platforms such as Kubernetes and Docker.
- Excellent incident response, troubleshooting, and documentation practices.
- Experience building and operating managed object, file, and block storage services across AWS, GCP, or Azure.
- Preferred: hands-on experience with Ceph, GlusterFS, OpenEBS, Vast, or Lightbits; open-source storage contributions; and hybrid on-premises and cloud storage experience.
Benefits
- Industry-competitive pay, restricted stock units, health insurance with HDHP and PPO options, vision and dental coverage, and employer HSA contributions.
- Paid parental leave, paid life insurance, short- and long-term disability, Teladoc, and a 401(k) with a 100% match up to 4% of salary.
- Generous paid time off and holiday schedule, cell phone reimbursement, tuition reimbursement, Calm app subscription, MetLife Legal, and a company-paid commuter benefit of $300 per month.
Tech Stack
Categories
Site Reliability
About Crusoe
Crusoe builds and operates GPU-powered data centers and an AI cloud platform for enterprises running large-scale AI training and inference. It vertically integrates energy supply with compute, using stranded and renewable power to reduce emissions and costs, and sells capacity via cloud services and managed infrastructure. Founded in 2018 and headquartered in Denver, the company is privately held.
