2 months ago
Guangzhou, ChinaStaff+
Responsibilities
- Define and lead the chaos engineering strategy, roadmap, and operating model for Treasury Technology.
- Establish standards, best practices, guardrails, metrics, reporting, and evidence for chaos engineering and resilience maturity.
- Design, review, and execute controlled chaos experiments and fault-injection activities across infrastructure, platforms, applications, and service dependencies.
- Identify resilience weaknesses, single points of failure, recovery gaps, and operational risks before production incidents occur.
- Drive remediation with engineering and platform teams and embed resilience-by-design principles into delivery and operational practices.
- Run Gamdays to rehearse failure scenarios and validate runbooks, alerting, and on-call readiness.
- Partner with architecture, SRE, DevOps, infrastructure, security, and support teams on resilience priorities.
- Coach engineers in resilience thinking, experiment design, and chaos engineering practices.
- Communicate technical findings, resilience risks, and improvement priorities to senior stakeholders and decision makers.
Requirements
- University degree or above in Computer Science, Software Engineering, or a related discipline.
- Strong experience in a technical leadership role in platform engineering, site reliability engineering, cloud engineering, resilience engineering, or a related discipline.
- Proven experience designing and implementing chaos engineering, fault injection, or resilience testing practices in complex enterprise environments.
- Deep hands-on knowledge of Kubernetes, including deployment behavior, failure handling, scaling, networking, and troubleshooting.
- Strong experience with GCP and cloud-native platforms, including operational and resilience considerations.
- Strong understanding of distributed systems, failure modes, system recovery, and service resilience patterns.
- Experience with modern technology stacks including microservices, APIs, platform services, and cloud-based applications.
- Strong understanding of observability, including metrics, logging, tracing, alerting, and service health monitoring.
- Experience with automation, CI/CD pipelines, and engineering tooling for scalable resilience practices.
- Strong problem-solving, stakeholder management, written communication, spoken communication, and influencing skills.
- Demonstrated resilience mindset focused on proactive risk reduction and continuous improvement.
- Excellent written and spoken English communication skills.
Benefits
- Flexible working.
- Opportunities for continuous professional development and career growth.
- Inclusive and diverse workplace committed to equal opportunity and respect.
Tech Stack
Categories
Site Reliability
About HSBC
Opening up a world of opportunity for our customers, investors, ourselves and the planet. We're a financial services organisation that serves more than 40 million customers, ranging from individual savers and investors to some of the world’s biggest companies and governments. Our network covers 58 countries and territories, and we’re here to use our unique expertise, capabilities, breadth and perspectives to open up a world of opportunity for our customers. HSBC is listed on the London, Hong Kong, New York, and Bermuda stock exchanges. To view our social media terms and conditions please visit the following webpage: http://www.hsbc.com/social-TandCs
