3 months ago
Base Salary
$209k - $253k/yr
Responsibilities
- Design and operate reliable managed AI services focused on serving and scaling LLM workloads
- Build automation and reliability tooling for distributed AI pipelines and inference services
- Define, measure, and improve SLIs and SLOs across AI workloads
- Collaborate with AI, platform, and infrastructure teams to optimize large-scale training and inference clusters
- Automate observability and develop telemetry and performance-tuning strategies for latency-sensitive AI services
- Investigate and resolve reliability issues in distributed AI systems using telemetry, logs, and profiling
- Contribute to the architecture of distributed systems designed for AI-first environments
Requirements
- Strong software engineering background building production-grade systems beyond scripting or Bash
- Demonstrated experience designing and implementing distributed systems
- Hands-on experience with large language models or AI/ML infrastructure
- Experience defining and measuring SLIs and SLOs, building monitoring and observability systems, and improving performance and reliability
- Experience designing fault-tolerant systems and automated testing strategies
- Proficiency in at least one of Python, Go, Java, or C++
- Familiarity with Kubernetes or container orchestration platforms
- Strong collaboration and communication skills
- Ability to thrive in a fast-paced, mission-driven environment
- Bonus: experience scaling LLM inference or training workloads
Benefits
- Restricted Stock Units in a fast-growing, well-funded technology company
- Health insurance options including HDHP and PPO, plus vision and dental coverage for dependents
- Employer HSA contributions
- Paid parental leave
- Paid life insurance and short- and long-term disability coverage
- Teladoc
- 401(k) with a 100% match up to 4% of salary
- Generous paid time off and holiday schedule
- Cell phone reimbursement
- Tuition reimbursement
- Calm app subscription
- MetLife Legal
- Company-paid commuter benefit
Tech Stack
Categories
About Crusoe
As the AI factory company, Crusoe is on a mission to accelerate the abundance of energy and intelligence. The company provides a reliable, scalable, cost-effective, and energy-first solution for AI infrastructure. By harnessing large-scale energy resources, building AI-optimized data centers, and delivering an AI cloud platform, Crusoe empowers its customers to build the future faster.
