Bybit

Senior Principal Site Reliability Engineer

Bybit
Apply
5 days ago
Hong Kong, Hong KongStaff+

Responsibilities

  • Design and build an enterprise-grade chaos engineering platform for multi-cluster, multi-region, and multi-environment Kubernetes and EC2 deployments.
  • Develop fault injection, production safety, traffic isolation, fault isolation, experiment orchestration, rollback, kill-switch, and real-time impact-monitoring capabilities.
  • Integrate chaos experiments with monitoring, alerting, and SLO systems to create an automated fault-injection and validation loop.
  • Define production fault-injection safety standards and approval workflows for mainnet systems.
  • Plan and execute daily, periodic, and large-scale cross-region or cross-AZ resilience and disaster-recovery experiments.
  • Create resilience scoring systems, recommend improvements, and drive remediation of identified weaknesses.
  • Evaluate Chaos Mesh, Litmus, custom components, and hybrid technology strategies.
  • Develop chaos engineering playbooks and mentor 2–3 engineers.

Requirements

  • 8+ years of backend or infrastructure engineering experience, including 3+ years dedicated to chaos engineering or stability engineering.
  • Hands-on experience with large-scale fault injection in production environments and a deep understanding of production safety constraints.
  • Expert proficiency in Kubernetes fault injection and familiarity with CRD and Operator development.
  • Proficiency in at least one backend language, preferably Go, with platform-level system architecture experience.
  • Deep understanding of distributed-system failure modes, including network partitions, split-brain, cascading failures, and data inconsistency.
  • Familiarity with Prometheus, Grafana, Thanos, and OpenTelemetry.
  • Strong technical documentation and solution-design skills.
  • Preferred experience includes financial or trading-system stability, SLO and Error Budget frameworks, automated fault recovery, AWS infrastructure, Chaos Mesh or Litmus contributions, and related chaos engineering tools.
  • Ability to balance production safety with validation depth, collaborate across teams, influence stakeholders, and work independently in ambiguous situations.

Benefits

  • Study Growth Fund supporting professional development and continuous learning.
  • Regular internal team-building events, workshops, and collaboration activities.
  • Global collaboration with colleagues across an international organization.
  • Career advancement opportunities within a rapidly expanding global company.
  • Internal mobility opportunities for long-term career development.

Tech Stack

AWSGoGrafanaKubernetesPrometheus

Categories

Site Reliability
Bybit

About Bybit

1,001-5,000 employees
Contact me