12 hours ago
Jakarta, IndonesiaSenior
Responsibilities
- Build automated validation, progressive rollout, and recovery mechanisms for safer production delivery.
- Define and operationalize SLOs, SLIs, and error budgets to guide reliability and delivery decisions.
- Create reusable AI-agent-driven workflows for alert triage, incident investigation, routine maintenance, and reporting.
- Improve logging, metrics, tracing, alerting, runbooks, incident investigation, and postmortems.
- Translate recurring incidents into tests, safeguards, automation, and lasting reliability improvements.
- Extend Terraform and delivery tooling to keep infrastructure reproducible and changes reviewable.
- Improve GCP infrastructure, networking, access controls, capacity management, reliability, and cost efficiency.
- Build maintainable operational tooling in Go, Python, or a comparable language and verify AI-generated work.
Requirements
- 6+ years of experience in SRE, production engineering, or infrastructure-heavy backend roles with hands-on production-system ownership.
- Strong production Google Cloud experience, including Cloud Run, networking, and IAM; Google Cloud expertise will be assessed during interviews.
- Fluency with Terraform and solid experience with deployment safety and delivery tooling.
- Strong observability, troubleshooting, root-cause analysis, and fix-verification skills.
- Programming ability in Go, Python, or a comparable language, with experience building maintainable operational tooling.
- Experience leading production incident investigations and implementing follow-up improvements.
- Practical fluency with AI coding agents and the ability to critically review, test, and validate their output.
- Strong written English and ability to collaborate asynchronously with a distributed team.
- Kubernetes or GKE experience, including deploying, operating, debugging, and scaling containerized services, is preferred.
- Datadog experience with dashboards, monitors, logs, APM, and distributed tracing is preferred.
- Experience with progressive delivery, policy-as-code, automated rollback, GCP cost management, or FinOps is preferred.
- Security experience is a plus.
Benefits
- Opportunity to shape reliability practices and automation for an AI-first engineering team.
- Work on systems with direct customer impact in enterprise Google Cloud environments.
- Small team with short decision paths, meaningful ownership, and appropriate review and production safeguards.
- AI-agent tools are part of the engineering workflow, with the opportunity to define how their impact is measured and improved.
- The company operates as a distributed team and emphasizes asynchronous collaboration.
Tech Stack
Categories
Site Reliability
About Aliz
Aliz provides cloud, data, and machine-learning consulting and implementation on Google Cloud for enterprises. The privately held firm, founded in 2009 and headquartered in Budapest, is a Premier Google Cloud partner with offices in Singapore, Berlin, and Jakarta. It builds scalable applications, data platforms, and analytics/ML pipelines, serving sectors such as airlines, retail, and telecom across Europe and Asia.
