
K8 Site Reliability SME
BitDeer Technologies Group11 days ago
Remote, United States or San Jose, CA, USASenior
Responsibilities
- Design, deploy, and operate production Kubernetes clusters optimized for GPU workloads at 100–10,000 GPUs.
- Configure Nvidia GPU Operator, device plugins, MIG, GPU time-slicing, topology-aware scheduling, NVLink locality, and network rail affinity.
- Develop CRDs and operators for GPU workload lifecycle management and enable automated drain, cordon, taint, and workload rescheduling around GPU failures.
- Integrate Slurm, Ray, and Kubeflow with Kubernetes.
- Implement multi-tenant isolation using namespaces, network policies, resource quotas, RBAC, and pod security standards.
- Automate bare-metal provisioning, tenant onboarding, lifecycle management, and reclamation through BMaaS.
- Build Terraform providers and modules for infrastructure-as-code across GPU clusters.
- Define and meet SLIs/SLOs for cluster availability, job completion, and provisioning latency.
- Operate monitoring, alerting, incident management, escalation, post-incident reviews, and executable runbooks.
- Enable safe AIOps-driven remediation and autonomous workflows in the Kubernetes control plane.
Requirements
- 5+ years of Kubernetes operations experience, including at least 2 years managing GPU workloads on Kubernetes.
- Deep knowledge of Nvidia GPU Operator, device plugins, GPU scheduling, topology-aware scheduling, and GPU-specific resource management.
- Hands-on experience building multi-tenant Kubernetes platforms with strong isolation guarantees.
- Experience with bare-metal server provisioning and lifecycle automation using Ironic, MAAS, or custom systems.
- Proficiency with Terraform, Helm, and GitOps workflows using ArgoCD or Flux.
- Strong SRE background including SLI/SLO frameworks, incident management, and capacity planning.
- Experience operating Prometheus, Grafana, and alerting at scale.
- Strong programming skills in Go or Python for operator and CRD development.
- Ability to design or implement autoscaler and remediator loops for Kubernetes and apply a runbook-as-code approach.
Tech Stack
Categories
DevOpsSite Reliability