4 hours ago
Base Salary
$215k - $260k/yr
Responsibilities
- Bring modern LLM inference optimization techniques into production.
- Design and optimize serving architectures, including prefill and decode disaggregation and request routing.
- Profile and improve serving systems from frameworks such as vLLM and SGLang down to CUDA kernels.
- Adapt optimization methods across ML models, especially large language models.
- Tune deployments for latency, throughput, cost, and reliability under real traffic.
- Partner with customer engineering teams to take workloads from proof of concept through monitored production deployment.
- Build and support production software and product features around the inference stack, primarily using Python.
- Define specifications, run focused experiments and proofs of concept, and ship tested optimizations.
- Own delivery from initial experimentation through production implementation and collaborate on features and product requirement documents.
- Make pragmatic tradeoffs around tooling and system complexity.
Requirements
- Bachelor's, Master's, or Ph.D. in Computer Science, Engineering, Mathematics, or a related field.
- Production software development experience with one or more general-purpose languages such as Python or C++, with Python preferred.
- Hands-on familiarity with techniques for optimizing LLMs for high-throughput and low-latency inference.
- Experience with LLM serving frameworks such as vLLM or SGLang and with kernel-level profiling and performance analysis.
- Strong understanding of GPU architecture and behavior.
- Clear interest in and hands-on experience with large language models.
- Working knowledge of AI/ML pipelines and the development and deployment lifecycle for ML models.
- Strong communication skills for explaining complex technical topics to customers and teammates.
- Preferred experience includes accelerating software systems, especially LLM systems, and building and shipping AI/ML inference systems.
- Preferred experience includes CUDA or comparable technologies, Docker, Kubernetes, and customer-facing AI/ML projects.
Benefits
- Competitive compensation and equity packages, including Restricted Stock Units.
- Paid time off, paid holidays, and leave of absence programs.
- Comprehensive health, dental, and vision insurance.
- Employer HSA contributions.
- Paid parental leave.
- Paid life insurance and short- and long-term disability coverage.
- Professional development and tuition reimbursement.
- Mental health and wellness support.
- Commuter benefits for parking and transit.
- Cell phone stipend.
- 401(k) retirement plan with company match up to 4% of salary.
- Volunteer time off.
- Global travel insurance and emergency assistance.
- Daily meals allowance and additional location-specific perks and programs.
Tech Stack
Categories
About Crusoe
As the AI factory company, Crusoe is on a mission to accelerate the abundance of energy and intelligence. The company provides a reliable, scalable, cost-effective, and energy-first solution for AI infrastructure. By harnessing large-scale energy resources, building AI-optimized data centers, and delivering an AI cloud platform, Crusoe empowers its customers to build the future faster.
