5 days ago
Kuala Lumpur, MalaysiaSenior
Responsibilities
- Develop and refactor high-concurrency, low-latency recommendation serving engines covering recall, coarse ranking, fine ranking, and re-ranking.
- Implement dynamic compute trimming, degradation mechanisms, personalization, and global strategy dispatch for traffic resilience.
- Build real-time feature streams and sliding-window aggregations using Kafka and Flink, and improve online/offline feature consistency in unified feature stores.
- Construct and optimize large-scale vector retrieval and heterogeneous indexing systems, targeting P99 retrieval latency below 100 milliseconds.
- Build embedding pipelines and inverted-index services for trading products, news, KOL content, on-chain signals, and other data sources.
- Deploy and optimize deep ranking models, including DIN, SIM, MMoE, and PLE, through quantization, graph optimization, and batching.
- Maintain recommendation-service availability above 99.9% and P99 latency below 200 milliseconds through overload protection, thread isolation, and disaster recovery mechanisms.
- Build distributed tracing and monitoring systems and contribute to A/B experimentation infrastructure, CUPED variance reduction, and sequential testing.
Requirements
- 5+ years of recommendation-system engineering experience at consumer-scale internet companies.
- Experience building or leading real-time recommendation systems serving tens of millions of users and preferably delivering systems from zero to one.
- Strong low-level computer science fundamentals and proficiency in at least one of Go, Java, or C++, with Go preferred and C++ beneficial.
- Experience with PyTorch or TensorFlow model inference and deployment optimization.
- Hands-on experience with Spark, Flink, and Kafka, including solving stream-computing latency and data-backlog problems.
- Proficiency in Milvus or Faiss cluster deployment and tuning.
- Deep understanding of collaborative filtering, two-tower retrieval, multi-objective optimization, MMoE, and PLE.
- Experience designing and developing recommendation, feature, or experimentation platforms or high-performance RPC frameworks.
Benefits
- Study Growth Fund supporting professional development and continuous learning.
- Internal team-building events, workshops, and collaboration activities.
- Global collaboration with an international team.
- Career advancement opportunities and internal mobility within a rapidly expanding company.
