
AI Research Engineer (Model Compression & Quantization)
Tether Operations Limited4 months ago
Remote, WorldwideSenior
Responsibilities
- Apply low-bit and mixed-precision quantization, including Quantization-Aware Training and Post-Training Quantization, to LLMs, VLMs, and other multimodal generative models.
- Use knowledge distillation to transfer capabilities from large teacher models to smaller student models across text, image, and audio inputs.
- Implement pruning methods to remove redundant parameters and attention heads while preserving task performance.
- Build compression pipelines and establish metrics for model size, latency, throughput, memory use, accuracy, and output fidelity.
- Analyze efficiency-versus-accuracy trade-offs and propose improvements based on empirical results.
- Research advanced compression methods such as adaptive pruning schedules and distillation with intermediate feature matching.
- Identify and address production inference bottlenecks for low-memory, low-latency edge deployment.
- Stay current with model-compression research for multimodal and generative architectures.
- Document methodologies, experiments, and results to support reproducibility and collaboration.
- Author technical papers and publish findings at leading conferences such as NeurIPS, ICML, ICLR, CVPR, ACL, and AAAI.
Requirements
- Bachelor’s degree in Computer Science or a related field.
- Ideally, a PhD in NLP, Machine Learning, or a related field, with a strong AI R&D track record and publications in top-tier conferences.
- Experience with PyTorch or an equivalent deep learning framework.
- Hands-on experience with model quantization, including Quantization-Aware Training and Post-Training Quantization.
- Research and hands-on experience with knowledge distillation for compressing large models.
- Research and hands-on experience with model pruning for compressing large models.
- Strong understanding of neural network architectures and training, including transformers, LLMs, VLMs, backpropagation, optimization, and fine-tuning.
- Familiarity with C++ is a plus, particularly for low-level quantization kernels or inference optimizations.
Benefits
- Remote work with a globally distributed team.
- Opportunity to work on advanced multimodal AI and model-compression research in a fintech and digital-asset environment.