Aliz

Senior Site Reliability Engineer

Aliz
Apply
12 hours ago
Jakarta, IndonesiaSenior

Responsibilities

  • Build automated validation, progressive rollout, and recovery mechanisms for safer production delivery.
  • Define and operationalize SLOs, SLIs, and error budgets to guide reliability and delivery decisions.
  • Create reusable AI-agent-driven workflows for alert triage, incident investigation, routine maintenance, and reporting.
  • Improve logging, metrics, tracing, alerting, runbooks, incident investigation, and postmortems.
  • Translate recurring incidents into tests, safeguards, automation, and lasting reliability improvements.
  • Extend Terraform and delivery tooling to keep infrastructure reproducible and changes reviewable.
  • Improve GCP infrastructure, networking, access controls, capacity management, reliability, and cost efficiency.
  • Build maintainable operational tooling in Go, Python, or a comparable language and verify AI-generated work.

Requirements

  • 6+ years of experience in SRE, production engineering, or infrastructure-heavy backend roles with hands-on production-system ownership.
  • Strong production Google Cloud experience, including Cloud Run, networking, and IAM; Google Cloud expertise will be assessed during interviews.
  • Fluency with Terraform and solid experience with deployment safety and delivery tooling.
  • Strong observability, troubleshooting, root-cause analysis, and fix-verification skills.
  • Programming ability in Go, Python, or a comparable language, with experience building maintainable operational tooling.
  • Experience leading production incident investigations and implementing follow-up improvements.
  • Practical fluency with AI coding agents and the ability to critically review, test, and validate their output.
  • Strong written English and ability to collaborate asynchronously with a distributed team.
  • Kubernetes or GKE experience, including deploying, operating, debugging, and scaling containerized services, is preferred.
  • Datadog experience with dashboards, monitors, logs, APM, and distributed tracing is preferred.
  • Experience with progressive delivery, policy-as-code, automated rollback, GCP cost management, or FinOps is preferred.
  • Security experience is a plus.

Benefits

  • Opportunity to shape reliability practices and automation for an AI-first engineering team.
  • Work on systems with direct customer impact in enterprise Google Cloud environments.
  • Small team with short decision paths, meaningful ownership, and appropriate review and production safeguards.
  • AI-agent tools are part of the engineering workflow, with the opportunity to define how their impact is measured and improved.
  • The company operates as a distributed team and emphasizes asynchronous collaboration.

Tech Stack

DatadogGoGoogle BigQueryGoogle CloudKubernetesPythonTerraform

Categories

Site Reliability
Aliz

About Aliz

201-500 employees

Aliz provides cloud, data, and machine-learning consulting and implementation on Google Cloud for enterprises. The privately held firm, founded in 2009 and headquartered in Budapest, is a Premier Google Cloud partner with offices in Singapore, Berlin, and Jakarta. It builds scalable applications, data platforms, and analytics/ML pipelines, serving sectors such as airlines, retail, and telecom across Europe and Asia.

Contact me