10 hours ago
Base Salary
$215k - $265k/yr
Responsibilities
- Define and drive the technical roadmap for services orchestrating Ray clusters across cloud and on-premises environments.
- Lead the design and optimization of control-plane components for large-scale heterogeneous AI and ML workloads.
- Establish organization-wide standards for infrastructure reliability, scalability, and observability.
- Direct long-term strategy for GPU, TPU, and container-management integration.
- Lead architecture discussions, resolve technical debt, and promote engineering excellence.
- Mentor engineers and influence technical direction across engineering, ML, customer-facing, and open-source teams.
Requirements
- 5+ years of experience writing high-quality production code and leading complex distributed-systems projects.
- Proven experience designing and maintaining highly available, scalable, and secure cloud-native platforms using AWS, Azure, or GCP.
- Deep expertise in Kubernetes-based deployments and container orchestration at massive scale.
- Advanced knowledge of the Linux kernel, networking, and low-level operating-system foundations.
- Mastery of Go and Python, including the ability to establish coding standards and best practices.
- Demonstrated ability to mentor senior engineers, influence without direct authority, and navigate complex technical trade-offs.
Benefits
- Competitive salary and equity.
- Health, dental, and vision coverage, with many plans up to 99% employer-covered.
- Flexible time off, paid parental leave, and mental health support.
Categories
About Anyscale
Anyscale builds a cloud platform and tools to run Ray, the open-source framework for distributed Python and AI/ML workloads, enabling teams to scale data prep, training, and inference. It monetizes through a managed service, enterprise features, and support for Ray deployments. Founded in 2019 and headquartered in San Francisco, this privately held company is the commercial steward of Ray, widely used to power production AI systems.
