21 days ago
Base Salary
$200k - $240k/yr
Responsibilities
- Design, build, and scale services orchestrating Ray clusters across cloud and on-premise environments.
- Optimize control-plane components for large-scale distributed AI/ML workloads.
- Build scheduling and resource-management systems for heterogeneous compute clusters.
- Improve the reliability, performance, scalability, and observability of managed Ray workloads.
- Support accelerator integration, including GPUs and TPUs.
- Develop container image management and dependency-resolution features for distributed workloads.
- Participate in code reviews and design and architecture discussions.
- Provide on-call support and troubleshoot infrastructure issues with customer and field teams.
- Collaborate with distributed-systems and machine-learning experts on AI infrastructure.
Requirements
- Bachelor’s degree in Computer Science, Engineering, or equivalent practical experience.
- At least 3 years of experience writing high-quality production code.
- Hands-on experience building and maintaining highly available, scalable, and performant distributed systems.
- Expertise with AWS, Azure, GCP, and Kubernetes-based deployments.
- Deep understanding of cloud networking, security, and authentication mechanisms.
- Familiarity with observability stacks such as Prometheus and Grafana.
- Proficiency in Go and Python.
- Knowledge of low-level operating-system foundations, including the Linux kernel, file systems, and containers.
Benefits
- The role offers opportunities to contribute to open-source Ray and Anyscale’s proprietary platform.
- The position includes on-call support responsibilities.
- Anyscale is an equal opportunity employer and an E-Verify company.
Tech Stack
Categories
DevOpsSite Reliability
About Anyscale
Anyscale builds a cloud platform and tools to run Ray, the open-source framework for distributed Python and AI/ML workloads, enabling teams to scale data prep, training, and inference. It monetizes through a managed service, enterprise features, and support for Ray deployments. Founded in 2019 and headquartered in San Francisco, this privately held company is the commercial steward of Ray, widely used to power production AI systems.
