12 hours ago
Base Salary
$250k - $300k/yr
Responsibilities
- Bring modern LLM inference techniques into production and optimize serving architectures for latency, throughput, cost, and reliability.
- Profile and analyze serving frameworks and CUDA kernels to identify and resolve performance bottlenecks.
- Adapt and scale optimization methods across ML models, especially large language models.
- Partner with customer engineering teams to tailor deployments and move workloads from proof of concept to monitored production services.
- Build and support production inference software and product features using general-purpose programming languages, preferably Python.
- Define specifications, create proofs of concept, run experiments, and ship well-tested optimization results.
- Own delivery from initial experimentation through production deployment and collaborate on feature and product requirement documents.
- Make sound technical tradeoffs and avoid unnecessary complexity.
Requirements
- Bachelor's, Master's, or Ph.D. in Computer Science, Engineering, Mathematics, or a related field.
- Production experience shipping code in a general-purpose language such as Python or C++, with a strong preference for Python.
- Hands-on experience optimizing large language models for high-throughput and low-latency inference.
- Familiarity with LLM serving frameworks such as vLLM or SGLang and with kernel-level profiling and performance analysis.
- Strong understanding of GPU architecture and behavior.
- Hands-on interest and experience with large language models and AI/ML development and deployment pipelines.
- Strong communication skills for explaining complex technical topics to customers and teammates.
- Preferred qualifications include software performance optimization experience, CUDA or comparable technology experience, AI/ML inference system development, Docker and Kubernetes experience, and customer-facing AI/ML project work.
Benefits
- Competitive compensation, bonus, equity packages, and Restricted Stock Units.
- Paid time off, paid holidays, leave programs, and paid parental leave.
- Health, dental, vision, HSA contributions, life insurance, and short- and long-term disability coverage.
- Professional development, tuition reimbursement, mental health and wellness support, and volunteer time off.
- Commuter benefits, cell phone stipend, 401(k) plan with company match up to 4% of salary, daily meal allowance, and global travel insurance.
- Additional location-specific perks and programs.
Tech Stack
Categories
About Crusoe
Crusoe builds and operates GPU-powered data centers and an AI cloud platform for enterprises running large-scale AI training and inference. It vertically integrates energy supply with compute, using stranded and renewable power to reduce emissions and costs, and sells capacity via cloud services and managed infrastructure. Founded in 2018 and headquartered in Denver, the company is privately held.
