24 hours ago
Tokyo, JapanSenior
Responsibilities
- Design, build, and operate large-scale machine learning systems across search and recommendation, including retrieval, ranking, personalization, and real-time inference.
- Productionize state-of-the-art algorithms with Applied Researchers by building scalable, reliable, and maintainable systems.
- Build embedding-generation, vector-indexing, approximate-nearest-neighbor, distributed-retrieval, and online model-serving systems.
- Optimize GPU and accelerator inference for latency, throughput, memory utilization, and infrastructure cost.
- Develop serving architectures using batching, caching, model compression, quantization, compilation, and distributed inference.
- Build reusable machine learning platforms and pipelines for training, evaluation, deployment, monitoring, experimentation, and model lifecycle management.
- Improve availability, observability, and operational excellence through monitoring, profiling, automated testing, and failure diagnosis.
- Lead cross-functional technical projects and define roadmaps that improve relevance, engagement, system performance, and business impact.
Requirements
- Bachelor’s or master’s degree in Computer Science, Computer Engineering, Machine Learning, or a related technical field.
- 7–10 years of relevant industry experience building production machine learning or large-scale distributed systems.
- Strong software engineering skills and proficiency in one or more of Python, Java, Scala, C++, or CUDA.
- Experience building production systems for search, recommendation, advertising, or other large-scale machine learning applications.
- Hands-on experience with embedding-based retrieval, vector indexing, approximate-nearest-neighbor search, distributed search, or sparse and dense retrieval.
- Experience optimizing deep learning inference on GPUs or other accelerators, including profiling, batching, quantization, compilation, memory optimization, or distributed inference.
- Proficiency with machine learning frameworks and serving technologies such as PyTorch, TensorFlow, ONNX, TensorRT, or Triton.
- Experience with large-scale data-processing technologies such as Spark, Hadoop, or SQL.
- Strong understanding of distributed systems and production engineering, including scalability, concurrency, latency optimization, observability, fault tolerance, capacity planning, and cloud or containerized infrastructure.
- Demonstrated ability to lead complex cross-functional projects and translate emerging machine learning research into reliable production solutions.
Categories
About TCGplayer
TCGplayer operates an online marketplace and software tools for buying and selling trading card games and related collectibles, serving hobby shops and individual sellers. It generates revenue through marketplace fees, seller subscriptions, and fulfillment programs such as TCGplayer Direct. An eBay subsidiary, the company provides authentication and logistics services that help stores list inventory at scale and reach buyers across the U.S. and internationally.
