1 month ago
Santa Clara, CA, USAMid Level / Senior
Base Salary
$135k - $200k/yr
Responsibilities
- Operate and evolve production Kubernetes clusters, including bare-metal provisioning automation, highly available control planes, node lifecycle, GPU container runtime, networking, and storage.
- Build safe, repeatable GitOps-based delivery for platform services and user applications using Argo CD, Helm, and Kustomize.
- Develop shared multi-tenant capabilities for scheduling, resource isolation, storage, networking, access control, secrets, and observability.
- Improve CPU/GPU utilization and cost efficiency across the compute platform.
- Build and improve reusable distributed batch and workflow platforms for Spark data processing and GPU-based replay and simulation.
- Perform work in accordance with the company’s Quality Management System requirements and contribute to continuous improvement.
Requirements
- Bachelor’s, master’s, or doctoral degree in Computer Science or a related technical field, or equivalent practical experience.
- Hands-on experience operating production Kubernetes clusters, including node lifecycle, upgrades, and troubleshooting.
- Experience with GitOps and infrastructure-as-code.
- Experience with GPU or ML workload scheduling, queueing and priorities, fractional GPU sharing, autoscaling, or multi-tenant resource management.
- Strong ownership, learning ability, and demonstrated ability to drive projects end to end.
- Preferred experience with Ray or Kubeflow.
- Preferred experience with lakehouse technologies such as Delta Lake or Apache Iceberg.
- Preferred experience operating large-scale distributed data-processing and workflow systems, including Apache Spark and Argo Workflows or an equivalent orchestrator.
Tech Stack
Categories
Data EngineeringDevOps
