
AI Research Engineer (Pre-training - LLM & Multi-Modal)
Tether Operations Limited3 months ago
Remote, WorldwideSenior
Responsibilities
- Conduct large-scale foundational pre-training for LLM and multi-modal models using distributed systems and thousands of NVIDIA GPUs.
- Design, prototype, and scale architectures, tokenizers, and cross-modal alignment layers.
- Source, filter, and curate large textual and multi-modal datasets and build efficient data pipelines.
- Execute experiments, analyze results, and refine training methods for performance and token efficiency.
- Debug model efficiency, computational performance, and multi-modal alignment bottlenecks during long training runs.
- Advance distributed training systems for scalability and hardware efficiency.
Requirements
- Degree in Computer Science or a related field.
- Hands-on experience contributing to large-scale LLM or multi-modal pre-training runs on distributed servers with thousands of NVIDIA GPUs.
- Practical experience with large-scale distributed training frameworks, libraries, and tools.
- Deep knowledge of state-of-the-art transformer and non-transformer model modifications for intelligence, efficiency, and scalability.
- Strong expertise in PyTorch and Hugging Face libraries, including model development, continual pre-training, and deployment.
- Solid AI R&D track record; a PhD in NLP, Machine Learning, or a related field and strong A* conference publications are preferred.
Benefits
- Remote work with a globally distributed team.
Tech Stack
Categories
AI Research