3 months ago
Hanoi, Vietnam or Ho Chi Minh City, VietnamSenior
Responsibilities
- Monitor, triage, and resolve platform incidents within defined SLA windows.
- Provision, scale, and operate cloud infrastructure on AWS and/or GCP.
- Maintain and improve CI/CD pipelines and GitOps workflows.
- Operate production-scale monitoring, logging, alerting, and observability systems.
- Participate in the global follow-the-sun on-call rotation.
- Configure, deploy, and manage AI tooling and MCP servers in production.
- Develop infrastructure automation, scripts, and internal tooling.
- Write post-incident reviews and contribute to monthly operational reporting.
- Collaborate with engineering teams across multiple time zones.
Requirements
- 4+ years of experience in a DevOps, SRE, or Platform Engineering role within an international team.
- Solid Kubernetes knowledge, including cluster operations, troubleshooting, and configuration.
- Hands-on experience with AWS and/or GCP.
- Understanding of networking fundamentals including DNS, load balancing, firewalls, and VPCs.
- Scripting and automation skills in Python, Bash, or similar.
- Experience with CI/CD tools and GitOps-based delivery.
- Working knowledge of monitoring and observability systems such as Prometheus and ELK.
- Good English communication skills for daily collaboration with European stakeholders.
- Self-directed and proactive approach to driving issues through resolution.
- Preferred: experience managing MCP servers and AI tooling, AI enablement or LLM infrastructure, eCommerce or SaaS platform support, or the Frontastic/commercetools Frontend ecosystem.
Benefits
- Performance bonus of up to 2 months’ salary.
- Performance reviews twice a year.
- Premium healthcare and an annual health check.
- 15 days of annual leave.
- Full salary during probation.
- Hybrid working; candidates based in Hanoi or Can Tho may work remotely, while offices are available in Ho Chi Minh City and Da Nang.
- Monday–Friday working schedule from 9 AM to 6 PM.
- Monthly Happy Hour, Community Tech activities, training programs, and global innovation projects.
- Career growth, leadership development, and mentorship opportunities.
Tech Stack
Categories
DevOpsSite Reliability
