Bybit

Senior Principal Site Reliability Engineer

Bybit
Apply
5 days ago
Kuala Lumpur, MalaysiaStaff+

Responsibilities

  • Design and build a multi-cluster, multi-region, multi-environment chaos engineering platform.
  • Develop fault injection, production safety, traffic isolation, fault isolation, experiment orchestration, monitoring, and automated pass/fail capabilities.
  • Define mainnet fault-injection safety standards and approval workflows.
  • Plan and execute daily, periodic, and large-scale disaster-recovery resilience experiments.
  • Create resilience scoring systems, recommend improvements, and drive remediation with business teams.
  • Evaluate Chaos Mesh, Litmus, and custom components; develop best practices and playbooks.
  • Mentor and grow 2–3 engineers in chaos engineering capabilities.

Requirements

  • 8+ years of backend or infrastructure engineering experience, including 3+ years dedicated to chaos engineering or stability engineering.
  • Hands-on experience with large-scale production fault injection and a deep understanding of production safety constraints.
  • Expert Kubernetes fault-injection experience with Chaos Mesh, Litmus, or custom solutions, including CRD and Operator development.
  • Proficiency in at least one backend language, preferably Go, and platform-level system architecture design.
  • Deep understanding of distributed-system failure modes such as network partitions, split-brain, cascading failures, and data inconsistency.
  • Familiarity with Prometheus, Grafana, Thanos, and OpenTelemetry.
  • Strong technical documentation and solution-design skills.
  • Preferred qualifications include financial or trading-system stability experience, SLO or error-budget framework experience, automated fault recovery experience, AWS infrastructure experience, knowledge of Netflix Chaos Engineering, AWS Fault Injection Simulator, or Gremlin, and open-source contributions to Chaos Mesh, Litmus, or similar projects.

Benefits

  • Study Growth Fund supporting professional development and continuous learning.
  • Internal team-building events, workshops, and collaboration activities.
  • Global collaboration with an international team.
  • Career advancement opportunities within a rapidly expanding company.
  • Internal mobility opportunities for long-term career development.

Tech Stack

AWSGoGrafanaKubernetesPrometheus

Categories

Site Reliability
Bybit

About Bybit

1,001-5,000 employees
Contact me