
AI Research Engineer (Model Compression & Quantization)
Tether Operations Limited4 months ago
Remote, WorldwideSenior
Responsibilities
- Apply low-bit and mixed-precision quantization to LLMs, VLMs, and other multimodal generative AI models while preserving accuracy and output quality.
- Use knowledge distillation to transfer capabilities from large teacher models to smaller student models across text, image, and audio inputs.
- Implement pruning methods to remove redundant parameters and attention heads while maintaining task performance.
- Analyze trade-offs among model size, latency, memory usage, throughput, and accuracy, and propose empirically supported improvements.
- Develop compression pipelines and apply advanced methods such as adaptive pruning schedules and intermediate feature matching.
- Establish performance and fidelity metrics and address production inference bottlenecks for edge deployment.
- Track current research in model compression for multimodal and generative architectures.
- Document methodologies, experiments, and results to support reproducibility, collaboration, and stakeholder communication.
- Author technical papers and publish model-compression research in top-tier conferences.
Requirements
- Bachelor’s degree in Computer Science or a related field.
- Ideally, a PhD in NLP, Machine Learning, or a related field, with a strong AI R&D publication record at leading conferences.
- Hands-on experience with PyTorch or equivalent deep learning frameworks.
- Hands-on experience with Quantization-Aware Training and Post-Training Quantization.
- Research and practical experience applying knowledge distillation to compress large models.
- Research and practical experience applying model pruning to compress large models.
- Strong understanding of neural network architectures and training, including transformers, LLMs, VLMs, backpropagation, optimization, and fine-tuning.
- Familiarity with C++ is preferred, particularly for low-level quantization kernels or inference optimization.
Benefits
- Remote work with a globally distributed team.
- Opportunity to conduct applied AI research and publish findings in major technical conferences.