
Senior Site Reliability Engineer
Carousell Group11 hours ago
Kuala Lumpur, MalaysiaSenior
Responsibilities
- Manage SLIs, SLOs, error budgets, reliability targets, and the balance between reliability work and feature delivery.
- Own incident detection, response, mitigation, postmortems, troubleshooting, on-call operations, and MTTD/MTTR improvement.
- Improve observability through metrics, logs, traces, dashboards, alerting, and service instrumentation.
- Design and maintain canary releases, progressive rollouts, automated rollback, and production delivery processes.
- Automate repetitive operational work and own infrastructure-as-code provisioning, configuration, policy, and documentation.
- Perform capacity planning, performance tuning, growth forecasting, headroom modeling, infrastructure upgrades, and cloud cost optimization.
- Implement infrastructure security controls, including secrets management, access control, monitoring, and response.
- Conduct production readiness reviews and resilience testing.
- Operate the MCP gateway and internal AI tooling, including availability, access control, rate limiting, cost visibility, and usage monitoring.
- Apply AI-assisted automation to incident triage, log summarization, and runbook generation.
Requirements
- At least 5 years of experience operating production systems at scale in SRE, DevOps, or infrastructure engineering.
- Production Kubernetes experience covering deployment, upgrades, troubleshooting, and maintenance.
- Strong Linux fundamentals and performance-tuning experience across RHEL, CentOS, Debian, or Ubuntu, plus Docker and container networking.
- Hands-on Google Cloud Platform experience with Terraform or equivalent infrastructure-as-code and declarative provisioning.
- Experience managing secrets and credentials with HashiCorp Vault or an equivalent system, including policies, rotation, and least-privilege access.
- Experience with CI/CD pipelines using GitHub Actions or an equivalent, including progressive delivery and automated rollback.
- Bash experience and working proficiency in Go, Python, or a similar general-purpose language for building production tooling.
- Experience with Prometheus, Grafana, and Google Cloud Operations for instrumentation, dashboards, alerting, and tracing.
- Ability to work independently on large, complex projects with minimal guidance.
- Practical experience applying AI and LLM-based tooling to engineering or operational workflows.
- Familiarity with MCP, agent harnesses, tool-calling loops, context management, guardrails, and failure handling, or the fundamentals and willingness to learn them quickly.
Benefits
- Opportunity to help shape an AI-first engineering organization and influence both the technology and engineering practices.
- The posting does not state compensation, work arrangement, or other specific benefits.
Tech Stack
Categories
DevOpsSite Reliability
About Carousell Group
Carousell Group runs a multi-category secondhand marketplace for consumers, delivered via mobile apps and web, across Greater Southeast Asia. It operates brands including Carousell, Mudah.my, Cho Tot, REFASH, Laku6, and others, and monetizes through transaction fees, advertising (Carousell Media Group), and recommerce/financing services. Founded in 2012 and headquartered in Singapore, the company serves tens of millions of users across seven markets.