Riot Games

Principal Software Engineer - DevOps / Site Reliability Engineer

Riot Games
Apply
about 3 hours ago
Singapore, SingaporeStaff+
H1B Sponsor

Responsibilities

  • Own and improve the reliability, availability, scalability, performance, and operational health of the AI Efficiency platform and its deployed tools.
  • Design and maintain infrastructure, deployment systems, release automation, environment management, and infrastructure-as-code practices.
  • Establish production-readiness standards covering monitoring, alerting, ownership, documentation, rollback, and support plans.
  • Define and operationalize SLIs, SLOs, error budgets, reliability metrics, observability, dashboards, synthetic monitoring, and actionable alerts.
  • Establish incident-management and on-call practices, lead production issue resolution, and drive post-incident corrective actions.
  • Build automation, progressive delivery, canary releases, feature flags, health checks, rollback mechanisms, and controlled environment promotion.
  • Perform capacity planning, load testing, performance analysis, resource forecasting, resilience validation, backup, recovery, failover, and disaster-recovery planning.
  • Identify systemic risks across applications, cloud infrastructure, networks, databases, queues, caches, dependencies, and operational workflows.
  • Improve developer experience through self-service workflows, reusable infrastructure, development and test environments, deployment tooling, and documentation.
  • Partner with software, ML platform, infrastructure, security, IT, developer-platform, compliance, and other teams to embed reliability, security, and maintainability into system design.
  • Support model-serving and inference systems and evaluate AI-assisted workflows for anomaly investigation, log analysis, regression detection, and runbook automation.
  • Define guardrails, approvals, auditability, escalation paths, and quality controls for automated or agentic systems interacting with production.
  • Provide technical leadership, mentoring, architecture reviews, standards, and tooling that improve operational excellence across the team.

Requirements

  • Bachelor’s degree in Computer Science or a related field, or equivalent professional experience.
  • At least five years of experience in SRE, DevOps, infrastructure, platform, production, developer experience, or similar roles supporting production systems.
  • Strong programming and automation skills in Python, Go, JavaScript, TypeScript, or comparable languages.
  • Experience operating cloud-based production systems in AWS, GCP, Azure, or comparable environments.
  • Experience with CI/CD, release systems, deployment automation, environment management, observability, incident response, and distributed systems.
  • Experience with containerized environments and orchestration technologies such as ECS, Docker, or Kubernetes.
  • Experience with infrastructure-as-code and configuration-management tools such as Terraform, Pulumi, or CloudFormation.
  • Working knowledge of Linux, networking, DNS, load balancing, service discovery, authentication, secrets management, and cloud security fundamentals.
  • Ability to identify systemic risks, influence technical direction across teams, and communicate during high-pressure incidents.
  • Experience providing technical leadership, mentoring engineers, and establishing engineering standards.
  • Preferred experience with AI/ML platforms, inference and model-serving systems, GPU-backed workloads, data pipelines, developer platforms, progressive delivery, resilience testing, disaster recovery, distributed-system components, security hardening, and operational governance.
  • Familiarity with Playwright and with AI-assisted tools or agentic systems used in incident investigation, testing, code review, remediation, or production operations.

Benefits

  • Full relocation support
  • Comprehensive health insurance for the employee, spouse, and children
  • Open paid time off
  • Retirement benefits with company matching
  • Life insurance, parental leave, and short-term and long-term disability coverage
  • Play Fund
  • Matching donations of time and money to non-profits
Riot Games

About Riot Games

1,001-5,000 employees

Since 2006, Riot Games has stayed committed to changing the way video games are developed, published, and supported for players. From our first title, League of Legends, to 2020’s VALORANT; we have strived to evolve the community with growth in Esports, and expansion from games into entertainment. Players are the foundation of Riot's community and because of them, we’re able to reach new heights. Founded by Brandon Beck and Marc Merrill, Riot is headquartered in Los Angeles, California, and has 4,500+ Rioters in 20+ offices worldwide. Riot has been featured on numerous lists including Fortune’s “100 Best Companies to Work For,” “25 Best Companies to Work in Technology,” “100 Best Workplaces for Millennials,” and “50 Best Workplaces for Flexibility.” Riot Games recruiters will never ask for money or request sensitive information, and they'll always reach out from an @riotgames.com email address. You can learn more about Riot’s interview process here: https://www.riotgames.com/en/work-with-us/interviewing-at-riot/interview-process