Global Fashion Group SGP Services Pte Ltd

Senior Site Reliability Engineer

Global Fashion Group SGP Services Pte Ltd
Apply
2 months ago
Ho Chi Minh City, VietnamSenior

Responsibilities

  • Operate and improve AWS infrastructure, Kubernetes clusters, CI/CD pipelines, monitoring, alerting, and deployment reliability.
  • Troubleshoot production issues across applications, infrastructure, networking, databases, queues, cloud services, and external integrations.
  • Optimize Kubernetes resource usage through requests and limits tuning, HPA, VPA, KEDA, Cluster Autoscaler, and node right-sizing.
  • Analyze AWS usage and cost data across EKS, EC2, RDS, S3, NAT Gateway, and other services to identify waste and savings opportunities.
  • Build automation for cost anomaly detection, unused-resource cleanup, and rightsizing recommendations.
  • Develop AIOps capabilities for alert ingestion, incident correlation, investigation workflows, and operational integrations.
  • Create dashboards, alerts, and runbooks for cost, reliability, and operational health.
  • Partner with software engineers and product managers to turn SRE, FinOps, and incident-management needs into platform features.

Requirements

  • Strong hands-on experience with AWS services including EKS, EC2, RDS or Aurora, OpenSearch, S3, CloudWatch, ALB, NAT Gateway, IAM, and VPC.
  • Strong experience operating Kubernetes and optimizing resource usage.
  • Experience with Terraform and infrastructure-as-code workflows.
  • Experience with observability tools such as Datadog, Prometheus, Grafana, and CloudWatch.
  • Experience building automation with Python, Go, TypeScript, or similar languages.
  • Understanding of APIs, webhooks, background workers, and event-driven workflows.
  • Ability to debug production issues across application, infrastructure, database, and networking layers.
  • Good Linux fundamentals and CI/CD experience.
  • Clear written and verbal English communication, practical decision-making, and an ownership mindset.
  • Preferred experience with AWS FinOps tools including Cost Explorer API, Compute Optimizer, CUR, Budgets, Savings Plans, or Reserved Instances.
  • Preferred experience with platform or internal tooling using FastAPI, Next.js, PostgreSQL, Redis, Celery, or React.
  • Preferred experience integrating Datadog, PagerDuty, GitHub, Jira, Slack, or OpenSearch APIs.
  • Preferred experience with AI-assisted operations, LLM integrations, agent-based workflows, security hardening, IAM least privilege, secrets management, and compliance practices.

Benefits

  • Hybrid working environment with work-from-home setup allowance.
  • MacBook or laptop provided when starting.
  • Social insurance, medical insurance, and AON insurance.
  • 13th month salary.
  • 15 days of annual leave, 30 days of sick leave or mental health leave, and 1 day of occasion leave.
  • Gym membership support.
  • Career progression plan and support.
  • Office amenities including massage chairs, table tennis, and a video game room.
  • Quarterly team events, yearly company trip, and end-of-year party.
  • Opportunity to work with a global talent pool in a growing fashion e-commerce organization.

Tech Stack

AWSDatadogFastAPIGoGrafanaKubernetesNext.jsPostgreSQLPrometheusPythonReactRedisTerraformTypeScript

Categories

DevOpsSite Reliability
Contact me