2 months ago
Base Salary
$170k - $205k/yr
Responsibilities
- Design and operate reliable managed AI services focused on serving and scaling LLM workloads
- Define, measure, and improve SLIs and SLOs for performance and reliability
- Collaborate with AI, platform, and infrastructure teams to optimize large-scale training and inference clusters
- Build telemetry and performance-tuning strategies for latency-sensitive services
- Investigate and resolve reliability issues in distributed AI systems using telemetry, logs, and profiling
- Contribute to next-generation distributed-systems architecture for AI-first environments
Requirements
- Strong software engineering background building production-grade systems beyond scripting or Bash
- Demonstrated experience designing and implementing distributed systems
- Experience defining and measuring SLIs and SLOs
- Experience building monitoring and observability systems
- Experience driving performance and reliability improvements
- Experience designing fault-tolerant systems and automated testing strategies
- Proficiency in at least one of Python, Go, Java, or C++
- Familiarity with Kubernetes or container orchestration platforms
- Strong collaboration and communication skills
- Ability to thrive in a fast-paced, mission-driven environment
Benefits
- Restricted Stock Units
- Health insurance options including HDHP and PPO, vision, and dental coverage for employees and dependents
- Employer HSA contributions
- Paid parental leave
- Paid life insurance and short- and long-term disability insurance
- Teladoc
- 401(k) with a 100% match up to 4% of salary
- Generous paid time off and holiday schedule
- Cell phone reimbursement
- Tuition reimbursement
- Calm app subscription
- MetLife Legal
- Company-paid commuter benefit
About Crusoe
Crusoe builds and operates GPU-powered data centers and an AI cloud platform for enterprises running large-scale AI training and inference. It vertically integrates energy supply with compute, using stranded and renewable power to reduce emissions and costs, and sells capacity via cloud services and managed infrastructure. Founded in 2018 and headquartered in Denver, the company is privately held.
