
Senior Cluster Site Reliability Engineer
The Voleon Group5 months ago
Remote, United States or Berkeley, CA, USASenior
Base Salary
$205k - $235k/yr
Responsibilities
- Respond to cluster outages and operational issues, triage incidents, and resolve urgent problems.
- Maintain high cluster uptime and define and track SLAs for reliability.
- Identify recurring operational problems and engineer systemic solutions with engineering teams.
- Develop cluster health metrics, telemetry, observability, and custom monitoring mechanisms.
- Collaborate with software and research teams on fair cluster usage policies and enforcement mechanisms.
- Forecast cluster growth, select scale-up strategies, and optimize operations for cost and usability.
Requirements
- 5+ years of experience in SRE or DevOps roles, preferably as a senior engineer or technical lead.
- Knowledge of HPC or batch compute frameworks and/or machine learning training systems.
- Ability to develop moderately complex scripts and utilities in a common scripting language.
- Familiarity with infrastructure-as-code and configuration management tools.
- Experience with AWS or GCP cloud infrastructure.
- Familiarity with modern observability stacks.
- Experience with distributed storage technologies.
- Bachelor’s degree in computer science.
- Preferred: hands-on experience with HPC frameworks, Kubernetes-based job orchestrators, and distributed computing frameworks.
- Preferred: familiarity with ML frameworks, hybrid or on-premises environments, HPC containerization, HPC networking, and security/IAM foundations.
Benefits
- Competitive compensation and benefits package.
- Technology talks from company experts.
- Modern office environment.
- Daily catered lunches.
Tech Stack
AnsibleApache AirflowApache SparkAWSDockerGoogle Cloud PlatformGrafanaKubernetesMLflowPrometheusPythonPyTorchRubyTensorFlowTerraform
Categories
DevOpsSite Reliability