BitDeer Technologies Group

K8 Site Reliability SME

BitDeer Technologies Group
Apply
11 days ago
Remote, United States or San Jose, CA, USASenior

Responsibilities

  • Design, deploy, and operate production Kubernetes clusters optimized for GPU workloads at 100–10,000 GPUs.
  • Configure Nvidia GPU Operator, device plugins, MIG, GPU time-slicing, topology-aware scheduling, NVLink locality, and network rail affinity.
  • Develop CRDs and operators for GPU workload lifecycle management and enable automated drain, cordon, taint, and workload rescheduling around GPU failures.
  • Integrate Slurm, Ray, and Kubeflow with Kubernetes.
  • Implement multi-tenant isolation using namespaces, network policies, resource quotas, RBAC, and pod security standards.
  • Automate bare-metal provisioning, tenant onboarding, lifecycle management, and reclamation through BMaaS.
  • Build Terraform providers and modules for infrastructure-as-code across GPU clusters.
  • Define and meet SLIs/SLOs for cluster availability, job completion, and provisioning latency.
  • Operate monitoring, alerting, incident management, escalation, post-incident reviews, and executable runbooks.
  • Enable safe AIOps-driven remediation and autonomous workflows in the Kubernetes control plane.

Requirements

  • 5+ years of Kubernetes operations experience, including at least 2 years managing GPU workloads on Kubernetes.
  • Deep knowledge of Nvidia GPU Operator, device plugins, GPU scheduling, topology-aware scheduling, and GPU-specific resource management.
  • Hands-on experience building multi-tenant Kubernetes platforms with strong isolation guarantees.
  • Experience with bare-metal server provisioning and lifecycle automation using Ironic, MAAS, or custom systems.
  • Proficiency with Terraform, Helm, and GitOps workflows using ArgoCD or Flux.
  • Strong SRE background including SLI/SLO frameworks, incident management, and capacity planning.
  • Experience operating Prometheus, Grafana, and alerting at scale.
  • Strong programming skills in Go or Python for operator and CRD development.
  • Ability to design or implement autoscaler and remediator loops for Kubernetes and apply a runbook-as-code approach.

Tech Stack

GoGrafanaHelmKubernetesPrometheusPythonTerraform

Categories

DevOpsSite Reliability
BitDeer Technologies Group

About BitDeer Technologies Group

201-500 employees
Contact me