5 hours ago
Base Salary
$185k - $225k/yr
Responsibilities
- Bring modern LLM inference optimization techniques into production.
- Design and optimize serving architectures, including prefill and decode disaggregation and request routing.
- Profile and optimize serving frameworks such as vLLM and SGLang down to CUDA kernel level.
- Adapt and scale optimization methods across ML models, especially large language models.
- Tune deployments for latency, throughput, cost, and reliability under real traffic.
- Partner with customer engineering teams to move workloads from proof of concept to monitored production services.
- Build and support production software and product features around the inference stack using general-purpose programming languages.
- Run focused experiments and proofs of concept, define clear specifications, and ship well-tested results.
- Own delivery from initial experimentation through production deployment and collaborate on feature and product requirement documents.
- Make pragmatic technical and tooling tradeoffs in ambiguous environments.
Requirements
- Bachelor's, Master's, or Ph.D. in Computer Science, Engineering, Mathematics, or a related field.
- Production software development experience in one or more general-purpose languages, such as Python or C++, with Python preferred.
- Hands-on familiarity with methods for optimizing LLMs for high-throughput and low-latency inference.
- Experience with LLM serving frameworks such as vLLM or SGLang and performance analysis down to the kernel level.
- Understanding of GPU architecture and behavior.
- Clear interest in and hands-on experience with large language models.
- Working knowledge of AI/ML pipelines and the end-to-end development and deployment of ML models.
- Strong communication skills for explaining complex technical topics to customers and teammates.
- Track record of improving software-system performance, especially for large language models, preferred.
- Experience with CUDA or comparable technologies preferred.
- Strong software engineering fundamentals and experience building and shipping AI/ML inference systems preferred.
- Experience with Docker and Kubernetes preferred.
- Customer-facing AI/ML project experience preferred.
Benefits
- Competitive compensation and equity packages, including Restricted Stock Units.
- Paid time off, paid holidays, and leave of absence programs.
- Comprehensive health, dental, and vision insurance.
- Employer HSA contributions.
- Paid parental leave.
- Paid life insurance and short- and long-term disability coverage.
- Professional development and tuition reimbursement.
- Mental health and wellness support.
- Commuter benefits for parking and transit.
- Cell phone stipend.
- 401(k) retirement plan with company match up to 4% of salary.
- Volunteer time off.
- Global travel insurance and emergency assistance.
- Daily meals allowance and additional location-specific perks and programs.
Tech Stack
Categories
About Crusoe
As the AI factory company, Crusoe is on a mission to accelerate the abundance of energy and intelligence. The company provides a reliable, scalable, cost-effective, and energy-first solution for AI infrastructure. By harnessing large-scale energy resources, building AI-optimized data centers, and delivering an AI cloud platform, Crusoe empowers its customers to build the future faster.
