BitDeer Technologies Group

Sr. SRE & Automation Engineer (Customer facing)

BitDeer Technologies Group
Apply
13 hours ago
Singapore, SingaporeSenior

Responsibilities

  • Own end-to-end reliability for the customer-facing GPU cloud service, including availability, job completion, provisioning latency, and tenant experience.
  • Operate production Kubernetes clusters for GPU workloads at 100–10,000 GPU scale, including GPU operators, device plugins, MIG, time-slicing, and topology-aware scheduling.
  • Manage customer and tenant onboarding, quotas, isolation, RBAC, resource policies, offboarding, and reclamation.
  • Build and operate Bare-Metal-as-a-Service provisioning, tenant handoff, lifecycle, and reclamation automation.
  • Define and publish customer-facing SLIs, SLOs, SLAs, error budgets, capacity plans, and reliability priorities.
  • Automate incident detection, remediation, escalation, customer communication, status updates, and post-incident reviews.
  • Develop tenant-aware monitoring, dashboards, alerting, GPU node failure handling, workload rescheduling, and self-service observability.
  • Build infrastructure-as-code and GitOps workflows across GPU clusters using Terraform, Helm, and ArgoCD or Flux.
  • Create runbooks, CRDs, remediation actuators, and workflow-engine integrations that support automated and self-healing operations.
  • Partner with customer success and support to turn customer-reported issues into systemic reliability improvements.

Requirements

  • 5+ years of experience in SRE or cloud operations, including at least 2 years operating GPU workloads at scale.
  • Deep Kubernetes operations and GPU workload-management experience, including Nvidia GPU Operator, device plugins, MIG, time-slicing, and GPU scheduling.
  • Experience with topology-aware scheduling and GPU-specific resource management.
  • Hands-on experience building multi-tenant cloud platforms with strong isolation guarantees.
  • Experience operating customer-facing cloud services, defining and meeting SLAs/SLOs, and managing tenant incidents and communications.
  • Experience with bare-metal server provisioning and lifecycle automation using Ironic, MAAS, or custom systems.
  • Proficiency with Terraform, Helm, and GitOps workflows using ArgoCD or Flux.
  • Strong knowledge of SLI/SLO/SLA frameworks, error budgets, incident management, and capacity planning.
  • Experience with Prometheus, Grafana, and alerting at scale.
  • Strong programming skills in Go or Python for automation and operator development.
  • Ability to design control planes for automated remediation and follow a runbook-as-code approach.

Benefits

  • Inclusive and respectful workplace that values diverse perspectives and backgrounds.
  • Opportunity to contribute directly to digital asset, Bitcoin mining, AI cloud, and HPC infrastructure projects.
  • Autonomy, personal accountability, fast growth, and learning opportunities.
  • Training, mentoring, welfare benefits, and opportunities to develop new processes and systems.
  • Global, fast-growing startup environment with opportunities to network with industry pioneers.

Tech Stack

GoGrafanaHelmKubernetesPrometheusPythonTerraform

Categories

DevOpsSite Reliability
BitDeer Technologies Group

About BitDeer Technologies Group

201-500 employees

Bitdeer Technologies Group (NASDAQ: BTDR) builds and operates Bitcoin‑mining infrastructure and AI/HPC cloud services for miners and enterprise compute customers. The Singapore‑headquartered public company, founded in 2021, designs ASIC chips and manufactures mining rigs, and runs data centers across the United States, Bhutan, Norway, Canada, Malaysia, and Ethiopia with a diversified 3 GW energy portfolio.

Contact me