AVP - System Development Manager - Data Platform - LME
Hong Kong Exchanges and Clearing Limited1 month ago
Shenzhen, ChinaStaff+
Responsibilities
- Design, build, operate, and continuously improve major streaming, batch, storage, or serving components of the data platform.
- Operate Kafka clusters, including broker tuning, partition rebalancing, monitoring, disaster recovery, schema registry, and Kafka Connect.
- Build and optimize Spark batch jobs and Flink or Spark Structured Streaming pipelines and frameworks.
- Manage Iceberg tables and operate MinIO storage at scale, including compaction, snapshot expiration, lifecycle rules, tiering, and performance tuning.
- Deploy and operate Trino, StarRocks, and ClickHouse clusters, including connector configuration, resource management, sharding, replication, materialized views, and query monitoring.
- Build and maintain Airflow or Dagster DAGs and extend custom operators and sensors.
- Implement platform-wide monitoring, dashboards, and alerting with OpenTelemetry, Prometheus, Grafana, and Loki.
- Manage Kubernetes operations, Helm charts, operators, resource quotas, node affinity, and pod disruption budgets.
- Own ArgoCD application sets, Helm-based deployments, and GitOps promotion pipelines from development to production.
- Participate in on-call rotations and write post-mortems and runbooks.
- Mentor mid-level engineers through pairing, design discussions, and code reviews.
Requirements
- 6+ years of experience in data engineering, platform engineering, or backend infrastructure.
- Strong Kubernetes experience, including Helm chart authoring, RBAC, network policies, persistent storage, and operators.
- Solid Kafka experience covering topic design, consumer groups, offsets, lag monitoring, Kafka Connect, and schema registries such as Apicurio or Confluent.
- Solid Spark experience with the DataFrame/Dataset API, Spark SQL, performance tuning, and production troubleshooting.
- Working knowledge of Flink or Spark Structured Streaming for real-time pipelines.
- Practical Iceberg experience with table maintenance, time travel, and catalog integration.
- Hands-on Trino or Presto experience with connector configuration, query tuning, and resource groups.
- Experience with an OLAP engine such as StarRocks, ClickHouse, or Doris, including table design, ingestion pipelines, and query optimization.
- Proficiency in Python and either Scala or Java.
- Experience with CI/CD and GitOps using ArgoCD or Flux, Helm, and Docker.
- Experience with Airflow or Dagster for pipeline orchestration.
- Preferred experience with OpenShift SCC, Routes, ImageStreams, and BuildConfigs; dbt; Great Expectations, Soda, or Deequ; Kafka Streams or ksqlDB; and DataHub or Atlas.
Benefits
- Permanent employment with a standard 40-hour scheduled workweek.
- Location: Shenzhen, China.
- Participation in an on-call rotation is required.
Tech Stack
Apache AirflowApache FlinkApache KafkaApache SparkClickHousedbtDockerGrafanaHelmJavaKubernetesOpenShiftPrestoPrometheusPythonScala
Categories
Data EngineeringDevOps