Tether Operations Limited

AI Research Engineer (Model Compression & Quantization)

Tether Operations Limited
Apply
4 months ago
Remote, WorldwideSenior

Responsibilities

  • Apply low-bit quantization to LLMs, VLMs, and other multimodal generative AI models while preserving accuracy and output quality.
  • Use knowledge distillation to transfer capabilities from large teacher models to smaller student models across text, image, and audio tasks.
  • Implement pruning methods to remove redundant parameters and attention heads while reducing computational overhead.
  • Measure and analyze trade-offs among model size, latency, memory consumption, throughput, and accuracy.
  • Research and implement mixed-precision quantization, adaptive pruning schedules, intermediate feature matching, and other advanced compression techniques.
  • Develop robust compression pipelines and address production inference bottlenecks for resource-constrained edge devices.
  • Stay current with research on compression for multimodal and generative architectures.
  • Document methodologies, experiments, and results to support reproducibility and collaboration.
  • Author technical papers and publish model-compression research at top-tier conferences.

Requirements

  • Bachelor’s degree in Computer Science or a related field.
  • PhD in NLP, Machine Learning, or a related field is preferred.
  • Solid track record in AI research and development, preferably including publications at A* conferences.
  • Experience with PyTorch or an equivalent deep learning framework.
  • Hands-on experience with both Quantization-Aware Training and Post-Training Quantization.
  • Research and practical experience using knowledge distillation to compress large models.
  • Research and practical experience using model pruning to compress large models.
  • Strong understanding of neural network architectures and training, including transformers, LLMs, VLMs, backpropagation, optimization, and fine-tuning.
  • Familiarity with C++ is preferred, especially for low-level quantization kernels or inference optimization.

Benefits

  • Remote work from locations around the world.
  • Opportunity to work on advanced multimodal AI and model-compression research within Tether’s AI research team.
  • Candidates should apply through Tether’s official careers channels; recruitment communication uses official company emails and platforms.

Tech Stack

Categories

AI ResearchML Engineering
Tether Operations Limited

About Tether Operations Limited

201-500 employees
Contact me