
Senior Site Reliability Engineer
Carousell Group7 hours ago
Kuala Lumpur, MalaysiaSenior
Responsibilities
- Operate services against SLIs, SLOs, and error budgets and adjust reliability targets as systems evolve.
- Own incident detection, response, mitigation, blameless postmortems, and follow-through across the full technology stack.
- Mature on-call operations by tuning alerts, maintaining runbooks, and improving MTTD and MTTR.
- Improve observability through instrumentation, metrics, logs, traces, dashboards, and alerting.
- Design and maintain canary releases, progressive delivery, and automated rollback across environments.
- Eliminate repetitive operational toil through automation and maintain infrastructure as code, configuration, policies, and documentation.
- Perform capacity planning, performance tuning, growth forecasting, headroom modeling, and cloud cost optimization.
- Implement infrastructure security controls, including secrets and credential management, access control, monitoring, and response.
- Conduct production readiness reviews and resilience testing.
- Operate the MCP gateway and internal AI tooling, including availability, access control, rate limiting, cost visibility, and usage monitoring.
- Apply AI-assisted automation to incident triage, log summarization, and runbook generation.
Requirements
- At least 5 years of experience operating production systems at scale in SRE, DevOps, or infrastructure engineering.
- Production Kubernetes experience covering deployment, upgrades, troubleshooting, and maintenance.
- Strong Linux fundamentals and performance-tuning experience with RHEL, CentOS, Debian, or Ubuntu, plus Docker and container networking.
- Hands-on Google Cloud Platform experience and infrastructure-as-code experience with Terraform or an equivalent tool.
- Experience managing secrets and credentials with HashiCorp Vault or an equivalent, including policies, rotation, and least-privilege access.
- Experience building CI/CD delivery pipelines with GitHub Actions or an equivalent, including progressive delivery and automated rollback.
- Bash proficiency and working proficiency in Go, Python, or another general-purpose programming language for building tooling.
- Experience with Prometheus, Grafana, and Google Cloud Operations for monitoring, dashboards, alerting, and tracing.
- Ability to work independently on large, complex projects with minimal guidance.
- Practical experience applying AI- and LLM-based tooling to engineering or operational workflows.
- Working familiarity with MCP, agent harnesses, tool-calling loops, context management, guardrails, and failure handling, or the fundamentals and willingness to learn them quickly.
Benefits
- Opportunity to help shape an AI-first engineering organization and its agentic workflows.
- Opportunity to influence both the technology built and the way engineering teams work.
- Role is with Mudah, part of Carousell Group, a recommerce group operating across Greater Southeast Asia.
Tech Stack
Categories
DevOpsSite Reliability
About Carousell Group
Carousell Group runs a multi-category secondhand marketplace for consumers, delivered via mobile apps and web, across Greater Southeast Asia. It operates brands including Carousell, Mudah.my, Cho Tot, REFASH, Laku6, and others, and monetizes through transaction fees, advertising (Carousell Media Group), and recommerce/financing services. Founded in 2012 and headquartered in Singapore, the company serves tens of millions of users across seven markets.