
Cloud Service Security Platform DevOps & Maintenance
BitDeer Technologies Group13 days ago
Remote, United States or San Jose, CA, USASenior
Responsibilities
- Design, implement, and maintain CI/CD pipelines for software applications and machine learning models, including automated build, testing, deployment, and rollback.
- Build, optimize, and scale cloud-native infrastructure with Kubernetes and Docker, including GPU clusters for AI workloads and model inference.
- Own high-availability production architecture, disaster recovery, self-healing, capacity planning, and performance tuning.
- Automate reproducible infrastructure provisioning across multiple cloud environments using Terraform, Ansible, and Helm.
- Develop monitoring, logging, and alerting systems for infrastructure, applications, and AI model metrics using Prometheus, Grafana, and ELK/EFK tools.
- Build and improve the Internal Developer Platform and golden deployment paths for product, model, and data science teams.
- Establish security, governance, release-management, Zero Trust, secrets-management, and compliance practices.
- Lead complex incident response, troubleshooting, root-cause analysis, and preventative automation.
- Collaborate with R&D, Data Science, Security, and Business teams to improve engineering workflows and efficiency.
Requirements
- Bachelor’s degree or above in Computer Science, Engineering, or a related technical field.
- At least 5 years of hands-on experience in DevOps, Site Reliability Engineering, or cloud infrastructure roles.
- Expertise in Linux, TCP/IP, DNS, HTTP, load balancing, and VPCs.
- Deep production-level knowledge of Docker and Kubernetes, including cluster management.
- Experience designing and managing infrastructure on AWS, GCP, Azure, Alibaba Cloud, or comparable public or hybrid cloud platforms.
- Strong programming or scripting ability in at least one of Go, Python, or Shell.
- Practical knowledge of CI/CD, Infrastructure as Code, observability, and SRE principles.
- Preferred experience with MLOps, vLLM, TGI, Triton Inference Server, GPU clusters, and AI/ML workloads.
- Preferred experience with large-scale distributed or high-concurrency systems and Internal Developer Platforms.
- Preferred familiarity with Zero Trust architecture, DevSecOps, SOC2, and ISO27001 compliance frameworks.
- Preferred technical leadership, mentoring, or DevOps team management experience.
- Preferred experience integrating an LLM-driven code or configuration helper into a pipeline.
Tech Stack
Alibaba CloudAnsibleAWSAzureDockerGoGoogle Cloud PlatformGrafanaHelmKubernetesLinuxPrometheusPythonTerraform