Responsibilities
- Design and build the sandboxing platform, client library, and API surface for secure code execution.
- Ensure strong isolation, security, and reproducibility across user sessions and workloads.
- Optimize cold-start latency, memory footprint, and resource utilization at scale.
- Reduce error rates through debugging, monitoring, and proactive fixes.
- Partner with internal teams to understand platform needs, debug issues, and build supporting tooling.
- Respond to incidents and production issues, conduct root cause analyses, and implement preventive fixes.
- Help maintain the sandboxing product roadmap and balance immediate needs with long-term architecture.
- Lead architecture reviews and own projects end-to-end from design through deployment.
Requirements
- 4+ years of experience building high-performance systems software, including meaningful experience maintaining libraries, SDKs, or developer-facing APIs.
- Deep understanding of Linux internals, including process isolation, memory management, cgroups, and namespaces.
- Experience with containerization and virtualization technologies such as Docker, Firecracker, gVisor, QEMU, and Kata Containers.
- Proficiency in a systems programming language such as Go, Rust, or C/C++.
- Experience with developer-facing API design, error propagation, documentation, and library usability.
- Comfort working across infrastructure layers from kernel modules to orchestration frameworks such as Kubernetes.
- Strong debugging skills and the ability to manage performance and security tradeoffs in production systems.
- Nice-to-have experience includes infrastructure startup ownership, LLM agents or agent frameworks, secure multi-tenant workloads, snapshotting and restore techniques, open-source systems contributions, and production on-call or incident response.
About Scale AI
Scale’s mission is to develop reliable AI systems for the world’s most important decisions. We provide the high-quality data and full-stack technologies that power the world’s leading models, and help enterprises and governments build, deploy, and oversee AI applications that deliver real impact. The Scale Generative AI Platform allows customers to build, evaluate, and control advanced AI agents and applications that continuously improve. The Scale Data Engine provides the technology to collect, curate, and annotate high-quality datasets. Through our Scale Labs, we test models with rigorous benchmarks and novel research to ensure breakthroughs translate into systems people can trust. Scale powers the most advanced LLMs and generative models in the world through RLHF, data generation and model evaluation. We work with industry leaders like Meta, Cisco, DLA Piper, Mayo Clinic, Time Inc., the Government of Qatar, and U.S. government agencies including the Army and Air Force.
