
AI Research Engineer (Model Compression & Quantization)
Tether Operations Limited4 months ago
Remote, WorldwideSenior
Responsibilities
- Apply low-bit and mixed-precision quantization, including QAT and PTQ, to LLMs, VLMs, and other multimodal generative AI models.
- Use knowledge distillation to transfer capabilities from larger teacher models to smaller student models across text, image, and audio inputs.
- Implement pruning methods, including removal of redundant parameters and attention heads, to reduce computational overhead.
- Build robust model-compression pipelines and establish metrics for model efficiency, fidelity, and production inference performance.
- Analyze size, latency, memory, throughput, and accuracy trade-offs and propose improvements based on empirical results.
- Research emerging compression strategies such as adaptive pruning schedules and intermediate feature-matching distillation.
- Optimize multimodal AI systems for low-memory, low-latency deployment on edge devices.
- Document methodologies, experiments, and results to support reproducibility and internal communication.
- Author technical papers and publish model-compression research in top-tier conferences.
Requirements
- Bachelor’s degree in Computer Science or a related field.
- Ideally, a PhD in NLP, Machine Learning, or a related field, with a strong AI R&D publication record at leading conferences.
- Hands-on experience with PyTorch or equivalent deep learning frameworks.
- Hands-on experience with model quantization, including Quantization-Aware Training and Post-Training Quantization.
- Research and practical experience with knowledge distillation for model compression.
- Research and practical experience with model pruning for model compression.
- Strong understanding of neural network architectures and training, including transformers, LLMs, VLMs, backpropagation, optimization, and fine-tuning.
- Familiarity with C++ is a plus, particularly for low-level quantization kernels or inference optimization.
- Excellent English communication skills.
Benefits
- Remote work with a globally distributed team.
- Opportunity to work on advanced multimodal AI and model-compression research in a fintech and digital-asset company.
- Opportunity to publish findings in leading AI and machine-learning conferences.