5 days ago
Hong Kong, Hong KongStaff+
Responsibilities
- Design and build an enterprise-grade chaos engineering platform for multi-cluster, multi-region, and multi-environment Kubernetes and EC2 deployments.
- Develop fault injection, production safety, traffic isolation, fault isolation, experiment orchestration, rollback, kill-switch, and real-time impact-monitoring capabilities.
- Integrate chaos experiments with monitoring, alerting, and SLO systems to create an automated fault-injection and validation loop.
- Define production fault-injection safety standards and approval workflows for mainnet systems.
- Plan and execute daily, periodic, and large-scale cross-region or cross-AZ resilience and disaster-recovery experiments.
- Create resilience scoring systems, recommend improvements, and drive remediation of identified weaknesses.
- Evaluate Chaos Mesh, Litmus, custom components, and hybrid technology strategies.
- Develop chaos engineering playbooks and mentor 2–3 engineers.
Requirements
- 8+ years of backend or infrastructure engineering experience, including 3+ years dedicated to chaos engineering or stability engineering.
- Hands-on experience with large-scale fault injection in production environments and a deep understanding of production safety constraints.
- Expert proficiency in Kubernetes fault injection and familiarity with CRD and Operator development.
- Proficiency in at least one backend language, preferably Go, with platform-level system architecture experience.
- Deep understanding of distributed-system failure modes, including network partitions, split-brain, cascading failures, and data inconsistency.
- Familiarity with Prometheus, Grafana, Thanos, and OpenTelemetry.
- Strong technical documentation and solution-design skills.
- Preferred experience includes financial or trading-system stability, SLO and Error Budget frameworks, automated fault recovery, AWS infrastructure, Chaos Mesh or Litmus contributions, and related chaos engineering tools.
- Ability to balance production safety with validation depth, collaborate across teams, influence stakeholders, and work independently in ambiguous situations.
Benefits
- Study Growth Fund supporting professional development and continuous learning.
- Regular internal team-building events, workshops, and collaboration activities.
- Global collaboration with colleagues across an international organization.
- Career advancement opportunities within a rapidly expanding global company.
- Internal mobility opportunities for long-term career development.
Tech Stack
Categories
Site Reliability
