5 days ago
Kuala Lumpur, MalaysiaStaff+
Responsibilities
- Design and build a multi-cluster, multi-region, multi-environment chaos engineering platform.
- Develop fault injection, production safety, traffic isolation, fault isolation, experiment orchestration, monitoring, and automated pass/fail capabilities.
- Define mainnet fault-injection safety standards and approval workflows.
- Plan and execute daily, periodic, and large-scale disaster-recovery resilience experiments.
- Create resilience scoring systems, recommend improvements, and drive remediation with business teams.
- Evaluate Chaos Mesh, Litmus, and custom components; develop best practices and playbooks.
- Mentor and grow 2–3 engineers in chaos engineering capabilities.
Requirements
- 8+ years of backend or infrastructure engineering experience, including 3+ years dedicated to chaos engineering or stability engineering.
- Hands-on experience with large-scale production fault injection and a deep understanding of production safety constraints.
- Expert Kubernetes fault-injection experience with Chaos Mesh, Litmus, or custom solutions, including CRD and Operator development.
- Proficiency in at least one backend language, preferably Go, and platform-level system architecture design.
- Deep understanding of distributed-system failure modes such as network partitions, split-brain, cascading failures, and data inconsistency.
- Familiarity with Prometheus, Grafana, Thanos, and OpenTelemetry.
- Strong technical documentation and solution-design skills.
- Preferred qualifications include financial or trading-system stability experience, SLO or error-budget framework experience, automated fault recovery experience, AWS infrastructure experience, knowledge of Netflix Chaos Engineering, AWS Fault Injection Simulator, or Gremlin, and open-source contributions to Chaos Mesh, Litmus, or similar projects.
Benefits
- Study Growth Fund supporting professional development and continuous learning.
- Internal team-building events, workshops, and collaboration activities.
- Global collaboration with an international team.
- Career advancement opportunities within a rapidly expanding company.
- Internal mobility opportunities for long-term career development.
Tech Stack
Categories
Site Reliability
