Base Salary
$201k - $251k/yr
Responsibilities
- Build, profile, and optimize the ML training and inference framework.
- Collaborate with ML teams to accelerate research and development and enable next-generation model and data-curation work.
- Research and integrate state-of-the-art technologies to optimize ML systems.
Requirements
- Experience with multi-node LLM training and inference.
- Experience developing large-scale distributed machine learning systems.
- Strong software engineering skills and proficiency with tools and frameworks such as CUDA, PyTorch, Transformers, and FlashAttention.
- Strong interest in system optimization and effective written and verbal communication skills.
- Ability to work in a cross-functional team environment.
- Expertise in LLM post-training methods or next-generation LLM use cases, including instruction tuning, RLHF, tool use, reasoning, agents, or multimodal systems, is preferred.
Benefits
- Comprehensive health, dental, and vision coverage.
- Retirement benefits, learning and development stipend, generous paid time off, and potentially a commuter stipend.
- Full-time position located in San Francisco, New York, or Seattle.
- Eligible roles may include equity-based compensation.
- Candidates must wait 90 days before reconsideration for the same role.
Tech Stack
Categories
About Scale AI
Scale’s mission is to develop reliable AI systems for the world’s most important decisions. We provide the high-quality data and full-stack technologies that power the world’s leading models, and help enterprises and governments build, deploy, and oversee AI applications that deliver real impact. The Scale Generative AI Platform allows customers to build, evaluate, and control advanced AI agents and applications that continuously improve. The Scale Data Engine provides the technology to collect, curate, and annotate high-quality datasets. Through our Scale Labs, we test models with rigorous benchmarks and novel research to ensure breakthroughs translate into systems people can trust. Scale powers the most advanced LLMs and generative models in the world through RLHF, data generation and model evaluation. We work with industry leaders like Meta, Cisco, DLA Piper, Mayo Clinic, Time Inc., the Government of Qatar, and U.S. government agencies including the Army and Air Force.