Tether Operations Limited

AI Research Engineer (Model Compression & Quantization)

Tether Operations Limited
Apply
4 months ago
Remote, WorldwideSenior

Responsibilities

  • Apply low-bit quantization to LLMs, VLMs, and other multimodal generative AI models while preserving accuracy and output quality.
  • Use knowledge distillation to transfer capabilities from large teacher models to smaller multimodal student models.
  • Implement pruning methods to remove redundant parameters and attention heads without materially reducing task performance.
  • Measure trade-offs among model size, latency, memory usage, throughput, and accuracy, and propose improvements based on experiments.
  • Research and apply mixed-precision quantization, adaptive pruning schedules, intermediate feature matching, and other advanced compression strategies.
  • Build compression pipelines, establish performance and fidelity metrics, and address production inference bottlenecks for edge deployment.
  • Track current research in compression for multimodal and generative architectures.
  • Document methodologies, experiments, and results to support reproducibility and collaboration.
  • Author technical papers and publish model-compression research in leading conferences such as NeurIPS, ICML, ICLR, CVPR, ACL, and AAAI.

Requirements

  • Bachelor’s degree in Computer Science or a related field.
  • PhD in NLP, Machine Learning, or a related field is preferred.
  • Solid track record in AI research and development, preferably including publications at A* conferences.
  • Hands-on experience with PyTorch or an equivalent deep learning framework.
  • Hands-on experience with both Quantization-Aware Training and Post-Training Quantization.
  • Research and practical experience using knowledge distillation to compress large models.
  • Research and practical experience using model pruning to compress large models.
  • Strong understanding of neural network architectures and training, including transformers, LLMs, VLMs, backpropagation, optimization, and fine-tuning.
  • Familiarity with C++ is preferred, particularly for low-level quantization kernels or inference optimizations.
  • Excellent English communication skills.

Benefits

  • Remote work with a globally distributed team.
  • Opportunity to conduct publication-oriented research on multimodal AI model compression and efficient edge deployment.
  • Candidates should apply through Tether’s official careers channels; recruitment communications use official company emails and platforms, and the company does not request payments or financial details.

Tech Stack

Categories

AI Research
Tether Operations Limited

About Tether Operations Limited

201-500 employees
Contact me