3 hours ago
Base Salary
$170k - $205k/yr
Responsibilities
- Lead the design and implementation of core managed AI services, including fault-tolerant queues, model catalogs, and scheduling mechanisms.
- Architect and scale infrastructure to support millions of API requests per second across thousands of customers.
- Implement monitoring and alerting to support system health and 24/7 availability.
- Collaborate with product management, business strategy, and engineering teams on the AI platform roadmap.
- Influence long-term platform architecture and technical decisions.
- Prototype emerging technologies, iterate on new features, and contribute to open-source AI frameworks.
Requirements
- Advanced degree in Computer Science or Engineering.
- 4–5+ years of industry experience leading varied initiatives with a record of consistent success.
- Experience with distributed systems, cloud compute, storage, networking, databases, and early-stage project delivery.
- Experience with generative AI, including LLMs and multimodal systems, and familiarity with AI infrastructure such as training, inference, and ETL pipelines.
- Proficiency with container runtimes such as Kubernetes, microservices, REST APIs, gRPC, and the full software development lifecycle.
- Preferred proficiency in Go, Python, or Rust for production services.
- Preferred contributions to open-source AI projects such as vLLM.
- Preferred experience optimizing GPU systems and inference frameworks.
Benefits
- Competitive compensation, bonus, and equity packages including Restricted Stock Units.
- Paid time off, holidays, leave programs, parental leave, and volunteer time off.
- Health, dental, vision, HSA contributions, life insurance, disability coverage, and mental health and wellness support.
- Professional development and tuition reimbursement.
- 401(k) retirement plan with company match up to 4% of salary.
- Commuter benefits, cell phone stipend, daily meals allowance, global travel insurance, emergency assistance, and location-specific perks.
- On-site work in San Francisco, California, or Sunnyvale, California.
Tech Stack
Categories
About Crusoe
Crusoe builds and operates GPU-powered data centers and an AI cloud platform for enterprises running large-scale AI training and inference. It vertically integrates energy supply with compute, using stranded and renewable power to reduce emissions and costs, and sells capacity via cloud services and managed infrastructure. Founded in 2018 and headquartered in Denver, the company is privately held.
