Tether Operations Limited

AI Research Engineer (Model Compression & Quantization)

Tether Operations Limited
Apply
4 months ago
Remote, WorldwideSenior

Responsibilities

  • Apply low-bit quantization to LLMs, VLMs, and other multimodal generative AI models while maintaining accuracy and output quality.
  • Use knowledge distillation to transfer capabilities from large teacher models to smaller multimodal student models.
  • Implement pruning methods to remove redundant parameters and attention heads without sacrificing task performance.
  • Analyze model-size, latency, memory, throughput, and accuracy trade-offs and propose improvements based on empirical results.
  • Research and apply mixed-precision quantization, adaptive pruning schedules, intermediate-feature distillation, and other advanced compression methods.
  • Build compression pipelines, establish performance and fidelity metrics, and address production inference bottlenecks on edge devices.
  • Stay current with model-compression research for multimodal and generative architectures.
  • Document methodologies, experiments, and results to support reproducibility and collaboration.
  • Author technical papers and publish findings at conferences including NeurIPS, ICML, ICLR, CVPR, ACL, and AAAI.

Requirements

  • Degree in Computer Science or a related field.
  • Strong AI research and development background with a track record of publications at leading conferences; a PhD in NLP, Machine Learning, or a related field is preferred.
  • Hands-on experience with PyTorch or an equivalent deep learning framework.
  • Experience with both Quantization-Aware Training and Post-Training Quantization.
  • Research and hands-on experience with knowledge distillation for compressing large models.
  • Research and hands-on experience with model pruning for compressing large models.
  • Strong understanding of neural network architectures and training, including transformers, LLMs, VLMs, backpropagation, optimization, and fine-tuning.
  • Familiarity with C++ is a plus, particularly for low-level quantization kernels or inference optimization.

Benefits

  • Remote work from anywhere in the world.

Tech Stack

Categories

AI Research
Tether Operations Limited

About Tether Operations Limited

201-500 employees
Contact me