
Senior AI Platform Engineer
BitDeer Technologies Group2 months ago
Singapore, SingaporeSenior
Responsibilities
- Operate and harden Kubernetes-based MaaS production environments across CPU nodes, edge ingress, and regional GPU tiers.
- Define and own SLOs, alerting, dashboards, runbooks, and incident response for API availability, latency, error rate, capacity, and GPU health.
- Improve rollout safety using canaries, fallback, health-aware routing, maintenance mode, and fast rollback.
- Drive capacity planning for GPU utilization, burst traffic, quota and rate limits, cross-region latency, and customer growth.
- Automate repetitive operations with Helm, Argo CD, operators, scripts, and self-healing workflows.
- Partner with runtime and performance engineers to debug incidents from the public API edge through model workers.
Requirements
- 6+ years of experience in SRE, platform engineering, or infrastructure engineering for production cloud services.
- Deep Kubernetes experience, including Helm, Argo CD, CNI and ingress, secrets, storage, and workload scheduling.
- Hands-on experience operating GPU, AI infrastructure, or HPC workloads is strongly preferred.
- Strong observability skills with Prometheus, VictoriaMetrics, OpenTelemetry, logs, traces, and incident diagnosis.
- Proficiency with Go, Python, Bash, Linux networking, and production automation.
- Ability to design reliable systems with clear SLOs, operational ownership, and post-incident follow-through.
Benefits
- Inclusive workplace that values authenticity and diverse backgrounds.
- Opportunity to contribute directly to the future of the digital asset industry and work on new projects.
- Personal accountability, autonomy, fast growth, and learning opportunities.
- Training, mentoring, welfare benefits, and developmental opportunities.
- Global, fast-growing company environment with networking opportunities and open workspaces.
Tech Stack
Categories
DevOpsSite Reliability