
Site Reliability Engineer
Razer Inc.4 months ago
Chengdu, ChinaMid Level
Responsibilities
- Design, develop, and maintain Infrastructure as Code using Terraform and AWS CloudFormation.
- Implement and operate reliable, scalable AWS infrastructure across compute, containerization, databases, storage, messaging, networking, and load balancing services.
- Lead and participate in architecture reviews focused on reliability, scalability, security, performance, and infrastructure cost efficiency.
- Develop and manage monitoring, alerting, and logging solutions using tools such as CloudWatch, Prometheus, Grafana, and ELK.
- Perform incident management, postmortems, root cause analysis, continuous improvement, and on-call support.
- Collaborate with software engineering teams on CI/CD pipelines, deployment automation, release management, and machine learning model deployment lifecycles.
- Automate infrastructure operations and reduce manual toil using Python, Bash, Node.js, Ruby, and AI-powered workflow automation.
- Maintain and troubleshoot web servers, databases, firewalls, DNS, load balancers, and networking environments.
- Apply security standards through patching, hardening, secure access policies, and protection of AI training data.
- Monitor and maintain SLOs, SLAs, and error budgets, and handle assigned incidents and support tickets.
Requirements
- Bachelor’s degree in Computer Science, Software Engineering, Information Technology, or a related field.
- Minimum three years of experience in SRE, DevOps, cloud infrastructure, or system administration roles.
- Hands-on expertise with AWS Cloud Services, including EC2, Lambda, ECS, EKS, Auto Scaling, load balancers, VPC, Route 53, Security Groups, firewalls, RDS, ElastiCache, Athena, S3, SQS, and SES.
- Deep understanding of Infrastructure as Code tools such as Terraform and CloudFormation.
- Proficiency in at least one programming or scripting language, including Python, Node.js, Bash, or Ruby.
- Experience operating and troubleshooting Linux, Windows, and container-based environments.
- Strong understanding of distributed systems, cloud networking, routers, switches, firewalls, DNS, and HTTP/TLS.
- Experience implementing monitoring and alerting systems and working with incident management processes.
- Experience with zero-downtime, blue/green, or canary deployments.
- Familiarity with AWS cost optimization, resource right-sizing, multi-region and multi-account architectures, API gateways, and edge networking such as Akamai or CloudFront.
- Experience with monitoring, logging, AIOps, predictive alerting, anomaly detection, and AI-assisted incident analysis is relevant to the role.
Benefits
- Razer states that it is an equal opportunity employer committed to an inclusive, respectful, and fair workplace.
- Reasonable accommodations are provided where needed, including for disability or religious practices.
- The role includes on-call support and participation in incident rotations.
About Razer Inc.
Razer builds gaming hardware, software, and services for PC and console gamers and esports audiences. Its products include peripherals and Blade laptops, plus software like Synapse, Chroma RGB, and Cortex; services include Razer Gold virtual credits and Razer Fintech payments in Southeast Asia. Founded in 2005 and dual-headquartered in Irvine, California, and Singapore, Razer operates globally as a privately held company.