Responsibilities
- Build and maintain fault-tolerant, high-performance systems for serving LLMs and other models at scale.
- Build an internal platform for LLM capability discovery.
- Collaborate with researchers and engineers to integrate and optimize models for production and research use cases.
- Conduct architecture and design reviews focused on system design and scalability.
- Develop monitoring and observability solutions for system health and performance.
- Lead projects end-to-end from requirements gathering through implementation in a cross-functional environment.
Requirements
- At least 4 years of experience building large-scale, high-performance backend systems.
- Strong programming skills in one or more of Python, Go, Rust, or C++.
- Experience with LLM serving and routing fundamentals such as rate limiting, token streaming, load balancing, and budgets.
- Experience with LLM capabilities and concepts including reasoning, tool calling, and prompt templates.
- Experience with Docker, Kubernetes, or similar container and orchestration tools.
- Familiarity with AWS or GCP and infrastructure as code such as Terraform.
- Ability to solve complex problems and work independently in fast-moving environments.
- Preferred experience with modern LLM serving frameworks such as vLLM, SGLang, TensorRT-LLM, or text-generation-inference.
Categories
About Scale AI
Scale’s mission is to develop reliable AI systems for the world’s most important decisions. We provide the high-quality data and full-stack technologies that power the world’s leading models, and help enterprises and governments build, deploy, and oversee AI applications that deliver real impact. The Scale Generative AI Platform allows customers to build, evaluate, and control advanced AI agents and applications that continuously improve. The Scale Data Engine provides the technology to collect, curate, and annotate high-quality datasets. Through our Scale Labs, we test models with rigorous benchmarks and novel research to ensure breakthroughs translate into systems people can trust. Scale powers the most advanced LLMs and generative models in the world through RLHF, data generation and model evaluation. We work with industry leaders like Meta, Cisco, DLA Piper, Mayo Clinic, Time Inc., the Government of Qatar, and U.S. government agencies including the Army and Air Force.
