
DevOps (GitLab-based platform CICD 30*3 pipelines)
BitDeer Technologies Group13 days ago
Remote, United States or San Jose, CA, USASenior
Responsibilities
- Design, implement, and maintain end-to-end CI/CD pipelines for software applications and machine learning models, including automated build, testing, deployment, and rollback workflows.
- Build, optimize, and scale cloud-native infrastructure with Kubernetes and Docker, including specialized GPU clusters for AI workloads and model inferencing.
- Own high-availability production architecture, disaster recovery, self-healing, capacity planning, and performance tuning.
- Automate reproducible and auditable infrastructure provisioning across multiple cloud environments using Terraform, Ansible, and Helm.
- Develop monitoring, logging, alerting, and AI model telemetry systems using tools such as Prometheus, Grafana, and ELK/EFK.
- Build and improve the Internal Developer Platform and golden paths for product, model, and data-science teams.
- Establish system stability, security, Zero Trust access controls, secrets management, release governance, and SOC2 and ISO27001 compliance practices.
- Lead troubleshooting and root-cause analysis during complex anomalies and major incidents, and implement preventative automation.
- Collaborate with R&D, Data Science, Security, and Business teams to improve engineering workflows and efficiency.
- Mentor engineers or serve as a technical lead when needed.
Requirements
- Bachelor’s degree or higher in Computer Science, Engineering, or a related technical field.
- At least 5 years of hands-on experience in DevOps, Site Reliability Engineering, or cloud infrastructure roles.
- Expert knowledge of Linux operating systems and networking concepts including TCP/IP, DNS, HTTP, load balancing, and VPCs.
- Deep production-level expertise with Docker and Kubernetes, including cluster management and orchestration principles.
- Experience designing and managing infrastructure on public or hybrid cloud platforms, including multi-cloud and hybrid-cloud strategies.
- Strong programming or scripting ability in at least one major language such as Go, Python, or Shell.
- Practical understanding of CI/CD, Infrastructure as Code, observability, and Site Reliability Engineering principles.
- Preferred experience with MLOps, model-serving or inferencing frameworks such as vLLM, TGI, or Triton Inference Server, and GPU clusters.
- Preferred experience with large-scale distributed systems or high-concurrency environments.
- Preferred experience designing and building Internal Developer Platforms.
- Preferred familiarity with Zero Trust architecture, automated security testing, DevSecOps, SOC2, and ISO27001.
- Preferred technical leadership, mentoring, or DevOps team management experience.
- Experience or strong judgment regarding LLM-driven code or configuration helpers and automation-first incident management is valued.
Tech Stack
Alibaba CloudAnsibleAWSAzureDockerGoGoogle Cloud PlatformGrafanaHelmKubernetesLinuxPrometheusPythonTerraform