
Sr SRE & Automation Engineer (Customer Facing)
BitDeer Technologies Group11 days ago
Remote, United States or San Jose, CA, USASenior
Responsibilities
- Own end-to-end reliability, availability, job completion, provisioning latency, and tenant experience for the customer-facing GPU cloud service.
- Operate production Kubernetes clusters optimized for GPU workloads at 100–10,000 GPU scale.
- Manage NVIDIA GPU Operator, device plugins, MIG, GPU time-slicing, topology-aware scheduling, and multi-tenant GPU allocation.
- Build tenant onboarding, quota management, isolation, offboarding, and reclamation workflows.
- Automate Bare-Metal-as-a-Service provisioning, tenant handoff, lifecycle management, and reclamation.
- Define and publish customer-facing SLIs, SLOs, SLAs, and error-budget priorities.
- Lead incident detection, remediation, escalation, customer communication, post-incident reviews, and recovery.
- Build tenant-aware monitoring, observability, dashboards, alerting, runbooks, and self-service operational tools.
- Automate GPU node failure handling, including detection, draining, cordoning, tainting, and workload rescheduling.
- Develop infrastructure-as-code and GitOps workflows across GPU clusters.
- Partner with customer success and support to turn tenant issues into systemic reliability improvements.
- Make remediation workflows and runbooks executable by the AIOps control plane.
Requirements
- 5+ years of experience in SRE or cloud operations.
- At least 2 years of experience operating GPU workloads at scale.
- Deep understanding of Kubernetes operations and GPU workload management, including NVIDIA GPU Operator, device plugins, MIG, time-slicing, and GPU scheduling.
- Experience with topology-aware scheduling and GPU-specific resource management.
- Hands-on experience building multi-tenant cloud platforms with strong isolation guarantees.
- Experience operating customer-facing cloud services against SLAs and SLOs, including tenant incident handling and communication.
- Experience with bare-metal server provisioning and lifecycle automation using Ironic, MAAS, or custom systems.
- Proficiency with Terraform, Helm, and GitOps workflows using ArgoCD or Flux.
- Strong knowledge of SLI, SLO, SLA, error-budget, incident-management, and capacity-planning practices.
- Experience with Prometheus, Grafana, and alerting at scale.
- Strong programming skills in Go or Python for automation and operator development.
- Ability to design automated remediation workflows and executable runbooks.
Tech Stack
Categories
DevOpsSite Reliability