
Senior Site Reliability Engineer (Night Shift 5PM-2AM)
Razer Inc.20 days ago
Ho Chi Minh City, VietnamSenior
Responsibilities
- Design, implement, and maintain Infrastructure as Code solutions using Terraform and/or CloudFormation across multi-account AWS environments.
- Collaborate with developers, architects, DevOps teams, and security teams to build scalable, secure, observable, and compliant cloud infrastructure.
- Lead architecture design sessions focused on reliability, scalability, security, and performance.
- Implement and manage monitoring, alerting, and observability solutions and track system reliability KPIs.
- Drive incident response, including coordination, triage, resolution, documentation, post-incident reviews, and technical investigations.
- Provide on-call support, participate in incident rotations, and lead investigations during outages or service degradation.
- Automate manual workflows using Python, Node.js, Bash, Ruby, JSON, and YAML.
- Troubleshoot infrastructure across Linux, Windows, Docker, and cloud-native environments.
- Support CI/CD and IaC best practices, disaster recovery, incident tickets, and solution handling.
- Supervise and mentor junior SREs and infrastructure engineers.
Requirements
- Bachelor’s degree in Computer Science, Software Engineering, Information Technology, or a related field.
- Minimum 3 years of experience in SRE, DevOps, cloud infrastructure, or system administration roles.
- Hands-on experience with AWS cloud services, including EC2, Lambda, ECS, EKS, Auto Scaling, Load Balancers, VPC, Route 53, Security Groups, Firewalls, RDS, ElastiCache, Athena, S3, SQS, and SES.
- Deep understanding of Terraform and CloudFormation for Infrastructure as Code.
- Proficiency in at least one of Python, Node.js, Bash, Ruby, or a related programming or scripting language.
- Experience operating and troubleshooting Linux, Windows, and container-based environments.
- Understanding of distributed systems, cloud networking, routers, switches, firewalls, DNS, HTTP, and TLS.
- Experience implementing monitoring and alerting systems and working with incident management processes.
- Experience with zero-downtime, blue/green, or canary deployments.
- Familiarity with AWS cost optimization, resource right-sizing, multi-region architecture, and multi-account architecture.
- Understanding of API gateways or edge networking technologies such as Akamai and CloudFront.
- Experience with disaster recovery, observability, incident response, and technical investigations.
Benefits
- Night-shift work schedule from 5:00 PM to 2:00 AM UTC+8 to provide continuous SRE coverage.
- Equal opportunity and inclusive workplace commitment across Razer’s global operations.
- Reasonable accommodations are available for disability and religious practices.
Categories
DevOpsSite Reliability