1 day ago
Base Salary
$230k - $295k/yr
Responsibilities
- Design and implement core services and abstractions for distributed compute infrastructure supporting thousands of concurrent jobs.
- Build a multiyear software engineering roadmap for the HPC platform with customer and infrastructure teams.
- Lead multi-quarter, cross-team initiatives that drive organization-wide improvements.
- Create production-grade APIs, SDKs, and tools for running large-scale distributed workloads.
- Design and improve job scheduling algorithms and auto-scaling policies for reliability and resource availability.
- Develop multi-region orchestration strategies optimized for data locality, reliability, and performance.
- Identify and resolve systemic reliability and performance issues through profiling, analysis, and collaboration with workload owners.
- Evaluate technologies and paradigms that improve computational and storage capabilities.
- Develop capacity-planning tools and forecasting models for growing compute needs.
- Mentor junior engineers and support their career development.
Requirements
- Experience designing and operating large-scale distributed systems in production.
- Experience with Ray.io, particularly Ray Core and Ray Data, or equivalent technologies.
- Experience with Kubernetes, particularly for heterogeneous workloads.
- Experience with cloud infrastructure on AWS or similar providers.
- Track record of shipping and operating reliable, scalable infrastructure.
- Ability to prioritize development work and build cross-functional consensus around technical tradeoffs.
- Proficiency with Python.
- Exposure to machine learning workloads such as training, inference, or data generation is a bonus.
- Experience with Kubernetes or SLURM at scale, including more than 10,000 nodes, is a bonus.
- Experience with SLURM workload management and advanced scheduling policies is a bonus.
- Background in algorithmic optimization or operations research is a bonus.
- Experience building developer tools and platforms used by large engineering organizations is a bonus.
Tech Stack
About Zoox
Zoox is transforming mobility-as-a-service by developing a fully autonomous, purpose-built fleet designed for AI to drive and humans to enjoy.