
AI Research Engineer (Model Compression & Quantization)
Tether Operations Limited4 months ago
Remote, WorldwideSenior
Responsibilities
- Apply low-bit quantization to LLMs, VLMs, and other multimodal generative AI models while maintaining accuracy and output quality.
- Use knowledge distillation to transfer capabilities from larger teacher models to smaller multimodal student models across text, image, and audio inputs.
- Implement pruning techniques to remove redundant parameters and attention heads without sacrificing task performance.
- Analyze efficiency-versus-accuracy trade-offs across quantization, distillation, and pruning methods and propose empirical improvements.
- Research and apply mixed-precision quantization, adaptive pruning schedules, intermediate feature matching, and other compression strategies.
- Build compression pipelines, establish performance and fidelity metrics, and address production inference bottlenecks.
- Document methodologies, experiments, and results to support reproducibility and collaboration.
- Author technical papers and publish model-compression research in top-tier AI conferences.
Requirements
- Bachelor’s degree in Computer Science or a related field.
- Solid track record in AI research and development, with strong publications in leading conferences; a PhD in NLP, Machine Learning, or a related field is preferred.
- Hands-on experience with PyTorch or an equivalent deep learning framework.
- Hands-on experience with Quantization-Aware Training and Post-Training Quantization.
- Research and hands-on experience with knowledge distillation for model compression.
- Research and hands-on experience with model pruning for model compression.
- Strong understanding of neural network architectures and training, including transformers, LLMs, VLMs, backpropagation, optimization, and fine-tuning.
- Familiarity with C++ is a plus, particularly for low-level quantization kernels or inference optimization.
Benefits
- Remote work with a globally distributed team.
- Opportunity to conduct applied AI research and publish findings in leading conferences.
- Work on efficient multimodal AI systems targeting resource-constrained edge devices.