BitDeer Technologies Group

Sr SRE & Automation Engineer (Customer Facing)

BitDeer Technologies Group
Apply
11 days ago
Remote, United States or San Jose, CA, USASenior

Responsibilities

  • Own end-to-end reliability, availability, job completion, provisioning latency, and tenant experience for the customer-facing GPU cloud service.
  • Operate production Kubernetes clusters optimized for GPU workloads at 100–10,000 GPU scale.
  • Manage NVIDIA GPU Operator, device plugins, MIG, GPU time-slicing, topology-aware scheduling, and multi-tenant GPU allocation.
  • Build tenant onboarding, quota management, isolation, offboarding, and reclamation workflows.
  • Automate Bare-Metal-as-a-Service provisioning, tenant handoff, lifecycle management, and reclamation.
  • Define and publish customer-facing SLIs, SLOs, SLAs, and error-budget priorities.
  • Lead incident detection, remediation, escalation, customer communication, post-incident reviews, and recovery.
  • Build tenant-aware monitoring, observability, dashboards, alerting, runbooks, and self-service operational tools.
  • Automate GPU node failure handling, including detection, draining, cordoning, tainting, and workload rescheduling.
  • Develop infrastructure-as-code and GitOps workflows across GPU clusters.
  • Partner with customer success and support to turn tenant issues into systemic reliability improvements.
  • Make remediation workflows and runbooks executable by the AIOps control plane.

Requirements

  • 5+ years of experience in SRE or cloud operations.
  • At least 2 years of experience operating GPU workloads at scale.
  • Deep understanding of Kubernetes operations and GPU workload management, including NVIDIA GPU Operator, device plugins, MIG, time-slicing, and GPU scheduling.
  • Experience with topology-aware scheduling and GPU-specific resource management.
  • Hands-on experience building multi-tenant cloud platforms with strong isolation guarantees.
  • Experience operating customer-facing cloud services against SLAs and SLOs, including tenant incident handling and communication.
  • Experience with bare-metal server provisioning and lifecycle automation using Ironic, MAAS, or custom systems.
  • Proficiency with Terraform, Helm, and GitOps workflows using ArgoCD or Flux.
  • Strong knowledge of SLI, SLO, SLA, error-budget, incident-management, and capacity-planning practices.
  • Experience with Prometheus, Grafana, and alerting at scale.
  • Strong programming skills in Go or Python for automation and operator development.
  • Ability to design automated remediation workflows and executable runbooks.

Tech Stack

GoGrafanaHelmKubernetesPrometheusPythonTerraform

Categories

DevOpsSite Reliability
BitDeer Technologies Group

About BitDeer Technologies Group

201-500 employees
Contact me