
DevOps Engineer III
Total System Services, LLC (TSYS)9 hours ago
Chengdu, ChinaSenior
Responsibilities
- Design, build, and maintain highly available, scalable, and resilient infrastructure and services across cloud and containerized environments.
- Drive reliability initiatives through chaos engineering, game days, incident analysis, performance testing, and disaster recovery validation.
- Develop automation and self-service capabilities using DevOps practices to improve operational efficiency and reduce manual effort.
- Build and enhance monitoring, alerting, and observability solutions to identify and resolve performance, security, and availability issues.
- Collaborate with Engineering and Technical Operations teams to troubleshoot infrastructure and network issues and improve platform stability.
- Contribute to reliability testing, incident management, continuous improvement, and operational excellence initiatives.
Requirements
- Bachelor's degree in Computer Science, Information Technology, Information Systems, or a related field.
- At least 2 years of experience in Site Reliability Engineering, DevOps, Cloud Infrastructure, Platform Engineering, or a related discipline.
- Experience supporting highly available production environments, including reliability, monitoring, incident management, and operational excellence.
- Experience with cloud platforms and container orchestration technologies such as Kubernetes, OpenShift, or AWS EKS.
- Hands-on experience with Terraform, Ansible, Jenkins, or similar infrastructure automation and configuration management tools.
- Strong troubleshooting and problem-solving skills across systems, networks, and application platforms.
- Knowledge of software delivery practices and modern DevOps methodologies.
- Preferred qualifications include chaos engineering experience, reliability testing leadership, GitOps and observability experience, and experience with large-scale distributed cloud-native systems.
Tech Stack
Categories
DevOpsSite Reliability