2 months ago
Palo Alto, CA, USASenior
Base Salary
$195k - $262k/yr
Responsibilities
- Own focused research projects from hypothesis through experiment, ablation, prototype, and production handoff.
- Prepare internal reports, technical blogs, or papers when the work is externally credible.
- Partner directly with MLEs to ensure research prototypes become usable production components.
- Define and execute research programs in efficient LLM and VLM inference with measurable production impact.
- Invent, evaluate, and productionize methods for quantization, QAT, distillation, and other optimization techniques.
- Build high-quality prototypes in PyTorch and related frameworks, then work with MLEs to productionize them.
- Design rigorous evaluation methodology covering various performance metrics.
- Publish papers, technical reports, and open-source artifacts to build external credibility.
- Collaborate with various teams to choose high-leverage research bets.
- Mentor engineers and scientists on experimental design and scientific rigor.
Requirements
- PhD in computer science, machine learning, or a closely related field.
- Strong publication record or equivalent research artifacts in relevant areas.
- Strong hands-on coding ability in Python and PyTorch.
- Deep understanding of LLMs, VLMs, and production-serving tradeoffs.
- Strong experimental design skills, including statistical reasoning.
- Excellent written and verbal communication skills.
Benefits
- 100% company-paid medical, dental, and vision coverage for employees and families.
- 401(k) plan with up to 4% company match and immediate vesting.
- 20 weeks paid parental leave for primary caregivers, 12 weeks for secondary caregivers.
- Remote work reimbursement of up to $85/month for mobile and internet.
- Company-paid short-term, long-term, and life insurance coverage.
About Nebius
The Nebius AI Cloud brings powerful full-stack infrastructure for AI developers and practitioners across startups, enterprises and science institutes to build and deploy generative AI applications and rapidly deliver scientific breakthroughs by training and running ML models within a secure, high-performance, and cost-optimized cloud environment.
