Senior Site Reliability Engineer
Global Fashion Group SGP Services Pte Ltd2 months ago
Ho Chi Minh City, VietnamSenior
Responsibilities
- Operate and improve AWS infrastructure, Kubernetes clusters, CI/CD pipelines, monitoring, alerting, and deployment reliability.
- Troubleshoot production issues across applications, infrastructure, networking, databases, queues, cloud services, and external integrations.
- Optimize Kubernetes resource usage through requests and limits tuning, HPA, VPA, KEDA, Cluster Autoscaler, and node right-sizing.
- Analyze AWS usage and cost data across EKS, EC2, RDS, S3, NAT Gateway, and other services to identify waste and savings opportunities.
- Build automation for cost anomaly detection, unused-resource cleanup, and rightsizing recommendations.
- Develop AIOps capabilities for alert ingestion, incident correlation, investigation workflows, and operational integrations.
- Create dashboards, alerts, and runbooks for cost, reliability, and operational health.
- Partner with software engineers and product managers to turn SRE, FinOps, and incident-management needs into platform features.
Requirements
- Strong hands-on experience with AWS services including EKS, EC2, RDS or Aurora, OpenSearch, S3, CloudWatch, ALB, NAT Gateway, IAM, and VPC.
- Strong experience operating Kubernetes and optimizing resource usage.
- Experience with Terraform and infrastructure-as-code workflows.
- Experience with observability tools such as Datadog, Prometheus, Grafana, and CloudWatch.
- Experience building automation with Python, Go, TypeScript, or similar languages.
- Understanding of APIs, webhooks, background workers, and event-driven workflows.
- Ability to debug production issues across application, infrastructure, database, and networking layers.
- Good Linux fundamentals and CI/CD experience.
- Clear written and verbal English communication, practical decision-making, and an ownership mindset.
- Preferred experience with AWS FinOps tools including Cost Explorer API, Compute Optimizer, CUR, Budgets, Savings Plans, or Reserved Instances.
- Preferred experience with platform or internal tooling using FastAPI, Next.js, PostgreSQL, Redis, Celery, or React.
- Preferred experience integrating Datadog, PagerDuty, GitHub, Jira, Slack, or OpenSearch APIs.
- Preferred experience with AI-assisted operations, LLM integrations, agent-based workflows, security hardening, IAM least privilege, secrets management, and compliance practices.
Benefits
- Hybrid working environment with work-from-home setup allowance.
- MacBook or laptop provided when starting.
- Social insurance, medical insurance, and AON insurance.
- 13th month salary.
- 15 days of annual leave, 30 days of sick leave or mental health leave, and 1 day of occasion leave.
- Gym membership support.
- Career progression plan and support.
- Office amenities including massage chairs, table tennis, and a video game room.
- Quarterly team events, yearly company trip, and end-of-year party.
- Opportunity to work with a global talent pool in a growing fashion e-commerce organization.
Tech Stack
Categories
DevOpsSite Reliability