27 days ago
Base Salary
$250k - $300k/yr
Responsibilities
- Bring modern inference techniques into production and refine them.
- Design and optimize serving architectures, including prefill/decode disaggregation and request routing.
- Profile and optimize serving frameworks such as vLLM and SGLang and the CUDA kernels underneath them.
- Adapt and scale optimization methods across ML models, especially large language models.
- Tune deployments for latency, throughput, cost, and reliability under real traffic.
- Partner with customer engineering teams to tailor deployments and move workloads from proof of concept to monitored production services.
- Build and support production software and product features around the inference stack.
- Turn ambiguous goals into specifications and proofs of concept, run experiments, and ship tested results.
- Own delivery from initial experimentation through production deployment and collaborate on features and product requirement documents.
- Make sound tradeoffs regarding complexity and tooling while maintaining accountability for performance goals and delivery.
Requirements
- Bachelor’s, Master’s, or Ph.D. in Computer Science, Engineering, Mathematics, or a related field.
- Production software development experience in one or more general-purpose languages, with Python preferred and C++ listed as an example.
- Experience with methods for optimizing LLMs for high-throughput and low-latency inference.
- Familiarity with vLLM or SGLang and with profiling and analyzing performance at the kernel level.
- Strong understanding of GPU architecture and behavior.
- Hands-on experience and clear interest in large language models.
- Working knowledge of AI/ML pipelines and the full model development and deployment lifecycle.
- Strong communication skills for explaining complex technical topics to customers and teammates.
- Preferred experience includes software performance optimization, CUDA or comparable technologies, software engineering fundamentals, AI/ML inference systems, Docker, Kubernetes, and customer-facing AI/ML work.
Benefits
- Competitive compensation and equity packages, including Restricted Stock Units
- Paid time off, paid holidays, and leave of absence programs
- Comprehensive health, dental, and vision insurance
- Employer HSA contributions
- Paid parental leave
- Paid life insurance and short- and long-term disability coverage
- Professional development and tuition reimbursement
- Mental health and wellness support
- Commuter benefits for parking and transit
- Cell phone stipend
- 401(k) retirement plan with company match up to 4% of salary
- Volunteer time off
- Global travel insurance and emergency assistance
- Daily meals allowance
- Additional location-specific perks and programs
Tech Stack
Categories
About Crusoe
As the AI factory company, Crusoe is on a mission to accelerate the abundance of energy and intelligence. The company provides a reliable, scalable, cost-effective, and energy-first solution for AI infrastructure. By harnessing large-scale energy resources, building AI-optimized data centers, and delivering an AI cloud platform, Crusoe empowers its customers to build the future faster.
