BitDeer Technologies Group

AI Cloud Senior DevOps Engineer

BitDeer Technologies Group
Apply
13 days ago
Remote, United States or San Jose, CA, USASenior

Responsibilities

  • Design, implement, and maintain end-to-end CI/CD and MLOps pipelines for applications and machine learning models.
  • Build, optimize, and scale Kubernetes- and Docker-based cloud-native infrastructure, including GPU clusters for AI workloads and model inferencing.
  • Own high-availability architecture, disaster recovery, self-healing, capacity planning, and production performance tuning.
  • Automate reproducible infrastructure provisioning across cloud environments using Terraform, Ansible, Helm, and infrastructure-as-code practices.
  • Develop monitoring, logging, and alerting systems for infrastructure, applications, and AI model metrics.
  • Partner with R&D, Data Science, Security, and Business teams on Internal Developer Platforms and platform engineering initiatives.
  • Establish security, governance, release management, secrets management, Zero Trust access controls, and compliance practices.
  • Lead major incident response, troubleshooting, root cause analysis, and preventative remediation.
  • Mentor engineers or manage DevOps teams when needed.

Requirements

  • Bachelor's degree or higher in Computer Science, Engineering, or a related technical field.
  • At least 5 years of hands-on experience in DevOps, Site Reliability Engineering, or cloud infrastructure roles.
  • Expert knowledge of Linux operating systems and networking concepts including TCP/IP, DNS, HTTP, load balancing, and VPCs.
  • Deep production-level expertise with Docker and Kubernetes, including cluster management and orchestration principles.
  • Experience designing and managing infrastructure on AWS, GCP, Azure, Alibaba Cloud, or other major public or hybrid cloud platforms.
  • Strong programming and scripting ability in at least one major language such as Go, Python, or Shell.
  • Practical understanding of CI/CD, infrastructure as code, observability, and Site Reliability Engineering principles.
  • Preferred experience with MLOps, vLLM, TGI, Triton Inference Server, GPU clusters, and AI/ML workloads.
  • Preferred experience with large-scale distributed systems, high-concurrency environments, Internal Developer Platforms, Zero Trust architecture, DevSecOps, SOC2, and ISO27001.
  • Preferred technical leadership, mentoring, or DevOps team management experience.

Tech Stack

Categories

BitDeer Technologies Group

About BitDeer Technologies Group

201-500 employees
Contact me