
AI Research Engineer (Model Compression & Quantization)
Tether Operations Limited4 months ago
Remote, WorldwideSenior
Responsibilities
- Apply low-bit quantization to LLMs, VLMs, and other multimodal generative AI models while preserving accuracy and output quality.
- Use knowledge distillation to transfer capabilities from large teacher models to smaller multimodal student models.
- Implement pruning methods to remove redundant parameters and attention heads without materially reducing task performance.
- Measure trade-offs among model size, latency, memory usage, throughput, and accuracy, and propose improvements based on experiments.
- Research and apply mixed-precision quantization, adaptive pruning schedules, intermediate feature matching, and other advanced compression strategies.
- Build compression pipelines, establish performance and fidelity metrics, and address production inference bottlenecks for edge deployment.
- Track current research in compression for multimodal and generative architectures.
- Document methodologies, experiments, and results to support reproducibility and collaboration.
- Author technical papers and publish model-compression research in leading conferences such as NeurIPS, ICML, ICLR, CVPR, ACL, and AAAI.
Requirements
- Bachelor’s degree in Computer Science or a related field.
- PhD in NLP, Machine Learning, or a related field is preferred.
- Solid track record in AI research and development, preferably including publications at A* conferences.
- Hands-on experience with PyTorch or an equivalent deep learning framework.
- Hands-on experience with both Quantization-Aware Training and Post-Training Quantization.
- Research and practical experience using knowledge distillation to compress large models.
- Research and practical experience using model pruning to compress large models.
- Strong understanding of neural network architectures and training, including transformers, LLMs, VLMs, backpropagation, optimization, and fine-tuning.
- Familiarity with C++ is preferred, particularly for low-level quantization kernels or inference optimizations.
- Excellent English communication skills.
Benefits
- Remote work with a globally distributed team.
- Opportunity to conduct publication-oriented research on multimodal AI model compression and efficient edge deployment.
- Candidates should apply through Tether’s official careers channels; recruitment communications use official company emails and platforms, and the company does not request payments or financial details.