9 days ago
Amsterdam, Netherlands or London, United KingdomSenior
Responsibilities
- Develop large-scale distributed training pipelines for complex models and datasets.
- Build and optimize low-latency inference pipelines for real-time production predictions.
- Develop libraries and scalable model frameworks to improve ML framework and system performance.
- Optimize training and inference using GPU hardware and acceleration libraries.
- Automate ML experiments, hyperparameter tuning, and model retraining with quantitative researchers.
- Partner with HPC specialists to improve training speed, workflow efficiency, and cost.
- Evaluate and deploy third-party tools for model development, training, and inference.
- Extend open-source ML tools by working with their internals and improving their capabilities.
Requirements
- 5+ years of machine learning experience focused on training or inference systems.
- Strong engineering skills in Python, CUDA, or C++.
- Knowledge of PyTorch, TensorFlow, or JAX.
- Proficiency in GPU programming and acceleration technologies such as CuDNN and TensorRT.
- Experience with distributed training technologies such as Horovod and NCCL.
- Experience with real-time, low-latency ML pipelines in high-performance environments is preferred.
- Exposure to cloud platforms and orchestration tools.
- Open-source contributions in machine learning, data science, or distributed systems are preferred.
Benefits
- Opportunity to join a brand-new position within a growing team.
