
Sr. Inference Service SRE & Automation Engineer
BitDeer Technologies Group7 hours ago
Singapore, SingaporeSenior
Responsibilities
- Own end-to-end availability, request success rate, latency, and SLA performance for customer inference endpoints.
- Operate and improve the API gateway, inference routing, scheduling, queueing, admission control, prefill/decode serving, and KV cache path.
- Manage inference runtimes, model loading and readiness validation, model rollouts, rollbacks, replica topology, and GPU fault isolation.
- Operate multi-cluster serving across Malaysia, Singapore, and Japan with traffic steering, failover, recovery, and evacuation procedures.
- Maintain GPU fleet health through monitoring, automated drain and cordon, failure diagnosis, and replica replacement.
- Operate the underlying Kubernetes platform and shared control-plane dependencies, including tested failover and decoupling serving from control-plane outages.
- Define SLIs, SLOs, SLAs, error budgets, burn-rate alerts, customer reporting, synthetic probes, and observability dashboards.
- Lead incident management, customer communication, postmortems, change safety, resilience testing, disaster recovery drills, and backup restoration verification.
- Drive capacity and saturation engineering, overload protection, load shedding, graceful degradation, and N+1 headroom.
- Build executable runbooks, CRDs, controllers, and workflow definitions for automated remediation and self-healing operations.
Requirements
- 5+ years of experience in SRE or production platform operations, including 2+ years operating a customer-facing API service against a published SLA.
- Deep production Kubernetes experience covering highly available control planes, etcd, CNI, service discovery, ingress and edge tiers, and multi-cluster topologies.
- Hands-on experience serving LLMs or other models with vLLM, SGLang, TensorRT-LLM, Triton, or equivalent runtimes.
- Experience with GPU fleet operations, including DCGM, Xid, ECC, NVLink, NCCL, node draining, and replica recovery.
- Experience with L4/L7 load balancing, gateways, active health checking, DNS/GSLB, and cross-cluster or cross-region failover.
- Strong SRE fundamentals covering SLI/SLO/SLA design, error budgets, incident command, postmortems, and capacity planning.
- Production experience operating PostgreSQL replication and failover, Redis, message queues, object storage, and secrets or certificate lifecycles.
- Observability experience with Prometheus- or VictoriaMetrics-class time-series databases, Grafana, log and trace pipelines, synthetic probing, and alert-quality management.
- Practical experience with admission control, concurrency and queue limits, load shedding, and graceful degradation.
- Proficiency with Terraform, Helm, and GitOps workflows using ArgoCD or Flux.
- Strong programming ability in Go or Python for automation, controllers, and operator development.
- AIOps aptitude, a runbook-as-code mindset, and comfort with customer-facing incident communication and SLA reporting.
Benefits
- Inclusive and respectful environment valuing diverse backgrounds and perspectives.
- Opportunity to work on new projects and directly influence digital asset and AI cloud infrastructure.
- Personal accountability, autonomy, fast growth, learning opportunities, training, mentoring, and developmental support.
- Welfare benefits and opportunities to network with industrial pioneers and technology enthusiasts.
- Workplace logistics and remote or hybrid arrangements are not specified.
Tech Stack
Categories
Site Reliability
About BitDeer Technologies Group
Bitdeer Technologies Group (NASDAQ: BTDR) builds and operates Bitcoin‑mining infrastructure and AI/HPC cloud services for miners and enterprise compute customers. The Singapore‑headquartered public company, founded in 2021, designs ASIC chips and manufactures mining rigs, and runs data centers across the United States, Bhutan, Norway, Canada, Malaysia, and Ethiopia with a diversified 3 GW energy portfolio.