
AI Cloud Senior DevOps Engineer
BitDeer Technologies Group13 days ago
Remote, United States or San Jose, CA, USASenior
Responsibilities
- Design, implement, and maintain end-to-end CI/CD and MLOps pipelines for applications and machine learning models.
- Build, optimize, and scale Kubernetes- and Docker-based cloud-native infrastructure, including GPU clusters for AI workloads and model inferencing.
- Own high-availability architecture, disaster recovery, self-healing, capacity planning, and production performance tuning.
- Automate reproducible infrastructure provisioning across cloud environments using Terraform, Ansible, Helm, and infrastructure-as-code practices.
- Develop monitoring, logging, and alerting systems for infrastructure, applications, and AI model metrics.
- Partner with R&D, Data Science, Security, and Business teams on Internal Developer Platforms and platform engineering initiatives.
- Establish security, governance, release management, secrets management, Zero Trust access controls, and compliance practices.
- Lead major incident response, troubleshooting, root cause analysis, and preventative remediation.
- Mentor engineers or manage DevOps teams when needed.
Requirements
- Bachelor's degree or higher in Computer Science, Engineering, or a related technical field.
- At least 5 years of hands-on experience in DevOps, Site Reliability Engineering, or cloud infrastructure roles.
- Expert knowledge of Linux operating systems and networking concepts including TCP/IP, DNS, HTTP, load balancing, and VPCs.
- Deep production-level expertise with Docker and Kubernetes, including cluster management and orchestration principles.
- Experience designing and managing infrastructure on AWS, GCP, Azure, Alibaba Cloud, or other major public or hybrid cloud platforms.
- Strong programming and scripting ability in at least one major language such as Go, Python, or Shell.
- Practical understanding of CI/CD, infrastructure as code, observability, and Site Reliability Engineering principles.
- Preferred experience with MLOps, vLLM, TGI, Triton Inference Server, GPU clusters, and AI/ML workloads.
- Preferred experience with large-scale distributed systems, high-concurrency environments, Internal Developer Platforms, Zero Trust architecture, DevSecOps, SOC2, and ISO27001.
- Preferred technical leadership, mentoring, or DevOps team management experience.
Tech Stack
Alibaba CloudAnsibleAWSAzureDockerGoGoogle Cloud PlatformGrafanaHelmKubernetesLinuxPrometheusPythonTerraform