2 months ago
Shenzhen, ChinaSenior
Responsibilities
- Support production infrastructure operations, release support, deployment automation, monitoring, troubleshooting, and recurring operational tasks.
- Maintain and improve infrastructure-as-code practices using Terraform and the team’s Git branching and review processes.
- Develop and maintain Shell and Python automation for deployment, updates, monitoring, and operational efficiency.
- Support AWS and Azure cloud operations, including infrastructure changes, access and configuration updates, reliability improvements, and cost and performance optimization.
- Operate and troubleshoot Docker and Kubernetes production environments, including deployment issues, workload health, scaling, rollbacks, and incident support.
- Participate in incident diagnosis, root cause analysis, post-incident follow-up, and preventive improvement actions.
- Improve observability using Prometheus, Grafana, and ELK through monitoring coverage, alert quality, dashboards, and log analysis.
- Prepare and maintain runbooks, SOPs, technical documentation, and handover materials.
- Coordinate with the DevOps and Infrastructure owner on priorities, architecture reviews, production changes, and stakeholder communication.
Requirements
- Bachelor’s degree in Computer Science, Engineering, or a related technical field, or equivalent practical experience.
- At least 5 years of experience in infrastructure, cloud operations, DevOps, distributed systems, or cloud migration.
- Strong Linux operations and troubleshooting skills with a solid understanding of networking fundamentals.
- Hands-on Shell and Python scripting experience for automation.
- Practical experience with AWS and/or Azure core services and cloud operations.
- Production experience with Terraform; multi-cloud experience is a plus.
- Production experience with Docker and Kubernetes, including troubleshooting, deployment, scaling, rollback, and operational support.
- Experience with CI/CD pipelines and application release automation.
- Familiarity with Prometheus, Grafana, ELK, or similar monitoring and logging platforms.
- Good documentation habits and willingness to work with runbooks, SOPs, tickets, and change records.
- Familiarity with Ansible, disaster recovery, resilience testing, or incident response automation.
