
Site Reliability Engineer II
Yum! Brands, Inc.1 month ago
Ho Chi Minh City, VietnamMid Level
Responsibilities
- Operate and improve reliable production systems through monitoring, observability, and reliability engineering practices.
- Participate in incident response, on-call or shift-based operations, and follow-the-sun regional handoffs.
- Automate manual operational work using scripting or programming languages.
- Work with Kubernetes, cloud infrastructure, infrastructure as code, and auto-healing or auto-remediation patterns.
- Contribute to platform engineering and internal developer platform initiatives, including self-service tooling and developer portals.
- Collaborate effectively with globally distributed teams through strong written and verbal English communication.
Requirements
- At least 2 years of experience in SRE, DevOps, production support, or infrastructure engineering roles.
- Hands-on experience with monitoring and observability tooling such as Datadog, Prometheus, Grafana, or CloudWatch.
- Working knowledge of at least one major cloud provider, with AWS preferred.
- Proficiency in at least one scripting or programming language such as Python, Bash, or Go, with demonstrated automation experience.
- Experience with incident response, on-call or shift-based operations, and SLI/SLO concepts.
- Experience with Kubernetes, container orchestration, and infrastructure as code such as Terraform.
- Familiarity with AI-assisted operations, auto-healing, auto-remediation, platform engineering, internal developer platforms, developer portals, and GitOps.
- Experience in multi-region or globally distributed team models is relevant.
- AWS, CKA, or similar certifications are relevant preferred qualifications.
Tech Stack
Categories
Site Reliability