
Site Reliability Engineer II
Yum! Brands, Inc.1 month ago
Ho Chi Minh City, VietnamMid Level
Responsibilities
- Monitor and improve production systems using observability tooling.
- Participate in incident response, on-call, and follow-the-sun shift operations.
- Automate manual operational work using scripting or programming languages.
- Develop reliability practices including auto-healing and auto-remediation.
- Contribute to Kubernetes, infrastructure as code, platform engineering, and internal developer platform initiatives.
- Coordinate structured handoffs and collaboration across globally distributed teams.
Requirements
- At least 2 years of experience in SRE, DevOps, production support, or infrastructure engineering.
- Hands-on experience with monitoring and observability tools such as Datadog, Prometheus, Grafana, or CloudWatch.
- Working knowledge of a major cloud provider, preferably AWS.
- Proficiency in Python, Bash, Go, or another scripting or programming language, with demonstrated automation experience.
- Experience with incident response, on-call or shift-based operations, and SLI/SLO concepts.
- Experience with Kubernetes, container orchestration, and infrastructure as code such as Terraform.
- Familiarity with AI-assisted operations, automation-first reliability, auto-healing, auto-remediation, platform engineering, developer portals, and globally distributed team models.
- Strong written and verbal English communication skills.
- Relevant certifications such as AWS or CKA are preferred.
Tech Stack
Categories
Site Reliability