29 days ago
Singapore, SingaporeSenior
Responsibilities
- Design and operate GitLab, AWS, and Kubernetes-based infrastructure for a government runtime platform.
- Develop automation and CI/CD pipelines to reduce toil and manual intervention.
- Implement observability using logs, metrics, traces, alerts, Golden Signals, SLOs, and Error Budgets.
- Participate in on-call rotations, respond to incidents, reduce MTTR, and conduct post-incident reviews.
- Design secure and compliant solutions in collaboration with security teams and support vulnerability scanning and audits.
- Identify performance bottlenecks, track reliability and cost KPIs, and drive platform optimization.
- Advise platform tenants on containerization and cloud-native deployment best practices.
- Create playbooks, runbooks, and documentation and share operational knowledge across teams.
- Stay current with AWS, Kubernetes, and industry developments and recommend platform improvements.
Requirements
- Bachelor's degree or Diploma in Computer Science, Engineering, or a related field, or equivalent experience.
- Proven experience as a Site Reliability Engineer or similar, with strong containerization, orchestration, and cloud-native technology experience.
- Experience troubleshooting complex technical issues in containerized applications and managing incidents, post-incident reviews, and continuous improvement.
- Deep understanding of AWS, Kubernetes, AWS EKS, networking, security, storage, and operational best practices, with familiarity with multi-cloud or hybrid environments.
- Experience integrating Kubernetes with AWS technologies such as Secrets Manager and Load Balancers and using Terraform or similar infrastructure-as-code tools.
- Hands-on experience with Kubernetes, Kustomize, Helm, and automation scripting using Go, Python, Bash, or equivalent.
- Ability to write and maintain automated tests or perform thorough manual testing for automation scripts.
- Familiarity with GitLab CI/CD, ArgoCD, and Git.
- Experience with Prometheus, Grafana, ELK Stack, observability, monitoring, SLOs, and Error Budgets.
- Certified Kubernetes Administrator or Certified Kubernetes Application Developer certification is a plus.
- Experience developing Kubernetes operators using Go, service mesh technologies, or Chaos Engineering is a plus.
- Strong problem-solving, analytical, communication, collaboration, documentation, customer advisory, mentoring, and incident-response skills.
Categories
Site Reliability
