2 months ago
Sydney, Australia or Hong Kong, Hong KongMid Level
Responsibilities
- Develop large-scale distributed training pipelines for datasets and complex models.
- Build and optimize low-latency inference pipelines for real-time production predictions.
- Develop libraries and scalable model frameworks to improve ML framework and model performance.
- Optimize training and inference using GPU hardware and acceleration libraries.
- Automate ML experiments, hyperparameter tuning, and model retraining with quantitative researchers.
- Partner with HPC specialists to improve training speed, workflow performance, and cost efficiency.
- Evaluate and deploy third-party tools for model development, training, and inference.
- Extend and improve open-source machine learning tools by working with their internals.
Requirements
- At least 3 years of machine learning experience focused on training or inference systems.
- Strong engineering skills in Python, CUDA, or C++.
- Knowledge of PyTorch, TensorFlow, or JAX.
- Proficiency in GPU programming and acceleration libraries such as CuDNN and TensorRT.
- Experience with distributed training tools such as Horovod and NCCL.
- Experience with real-time, low-latency ML pipelines in high-performance environments is a plus.
- Exposure to cloud platforms and orchestration tools is preferred.
- Contributions to open-source projects in machine learning, data science, or distributed systems are a plus.