BitDeer Technologies Group

Sr GPU Cloud K8S Expert (SRE SME)

BitDeer Technologies Group
Apply
3 hours ago
Remote, United States or San Jose, CA, USASenior

Responsibilities

  • Design, deploy, and operate production Kubernetes clusters optimized for GPU workloads at 100–10,000 GPUs.
  • Configure NVIDIA GPU Operator, device plugins, MIG, GPU time-slicing, topology-aware scheduling, GPU locality, NVLink awareness, and network rail affinity.
  • Develop and manage CRDs for GPU workload lifecycle management and integrate Slurm, Ray, and Kubeflow with Kubernetes.
  • Implement multi-tenant isolation using namespaces, network policies, resource quotas, RBAC, and pod security standards.
  • Automate bare-metal provisioning, tenant onboarding, lifecycle management, and reclamation for BMaaS.
  • Build Terraform providers and modules for GPU-cluster infrastructure automation.
  • Define and operate SLIs and SLOs for cluster availability, job completion rates, and provisioning latency.
  • Automate incident response, runbooks, escalation, post-incident reviews, GPU node failure detection, draining, cordoning, tainting, and workload rescheduling.
  • Operate monitoring and alerting with Prometheus, Grafana, Alertmanager, and PagerDuty.
  • Enable AIOps remediation workflows by making the Kubernetes control plane safe for autonomous action.

Requirements

  • 5+ years of experience in Kubernetes operations, including at least 2 years managing GPU workloads on Kubernetes.
  • Deep knowledge of NVIDIA GPU Operator, device plugins, and GPU scheduling in Kubernetes.
  • Experience with topology-aware scheduling, GPU-specific resource management, and multi-tenant Kubernetes platforms with strong isolation guarantees.
  • Hands-on experience with bare-metal server provisioning and lifecycle automation using Ironic, MAAS, or custom solutions.
  • Proficiency with Terraform, Helm, and GitOps workflows using Argo CD or Flux.
  • Strong SRE background including SLI/SLO frameworks, incident management, and capacity planning.
  • Experience operating Prometheus, Grafana, and alerting at scale.
  • Strong programming skills in Go or Python for operator and CRD development.
  • Ability to design or implement Kubernetes-integrated autoscaler or remediation loops.
  • Runbook-as-code mindset and ability to make SRE playbooks executable by the platform.

Tech Stack

Argo CDGoGrafanaHelmKubernetesPrometheusPythonTerraform

Categories

DevOpsSite Reliability
BitDeer Technologies Group

About BitDeer Technologies Group

201-500 employees

Bitdeer Technologies Group (NASDAQ: BTDR) builds and operates Bitcoin‑mining infrastructure and AI/HPC cloud services for miners and enterprise compute customers. The Singapore‑headquartered public company, founded in 2021, designs ASIC chips and manufactures mining rigs, and runs data centers across the United States, Bhutan, Norway, Canada, Malaysia, and Ethiopia with a diversified 3 GW energy portfolio.

Contact me