7 months ago
Responsibilities
- Design and implement improvements to model training infrastructure.
- Contribute to technical decisions that optimize model performance.
- Work on post-training processes, including reinforcement learning and fine-tuning.
- Improve the efficiency and scalability of model serving infrastructure.
- Profile GPU-accelerated workloads, identify bottlenecks, and implement performance optimizations.
Requirements
- Strong engineering skills with fluency in Python and PyTorch or other frameworks.
- Proven experience implementing and training large deep learning models.
- Experience writing and debugging low-level GPU code using CUDA and C++.
- Experience scaling GPU jobs with large-scale compute clusters such as Slurm or Kubernetes.
- Ability to analyze and optimize GPU-accelerated workloads through profiling, bottleneck identification, and performance tuning.
Benefits
- Remote-first work arrangement with a globally distributed team
- Five weeks of paid leave
- Comprehensive healthcare benefits including vision and dental
- Additional well-being perks
- Visa assistance, including H1B and OPT transfers, for US employees
Tech Stack
Categories
About Reka
AI research and product company with a mission to build models that understand the physical world.
