Tether Operations Limited

AI Research Engineer (Model Compression & Quantization)

Tether Operations Limited
Apply
4 months ago
Remote, WorldwideSenior

Responsibilities

  • Apply low-bit and mixed-precision quantization, including QAT and PTQ, to LLMs, VLMs, and other multimodal generative AI models.
  • Use knowledge distillation to transfer capabilities from larger teacher models to smaller student models across text, image, and audio inputs.
  • Implement pruning methods, including removal of redundant parameters and attention heads, to reduce computational overhead.
  • Build robust model-compression pipelines and establish metrics for model efficiency, fidelity, and production inference performance.
  • Analyze size, latency, memory, throughput, and accuracy trade-offs and propose improvements based on empirical results.
  • Research emerging compression strategies such as adaptive pruning schedules and intermediate feature-matching distillation.
  • Optimize multimodal AI systems for low-memory, low-latency deployment on edge devices.
  • Document methodologies, experiments, and results to support reproducibility and internal communication.
  • Author technical papers and publish model-compression research in top-tier conferences.

Requirements

  • Bachelor’s degree in Computer Science or a related field.
  • Ideally, a PhD in NLP, Machine Learning, or a related field, with a strong AI R&D publication record at leading conferences.
  • Hands-on experience with PyTorch or equivalent deep learning frameworks.
  • Hands-on experience with model quantization, including Quantization-Aware Training and Post-Training Quantization.
  • Research and practical experience with knowledge distillation for model compression.
  • Research and practical experience with model pruning for model compression.
  • Strong understanding of neural network architectures and training, including transformers, LLMs, VLMs, backpropagation, optimization, and fine-tuning.
  • Familiarity with C++ is a plus, particularly for low-level quantization kernels or inference optimization.
  • Excellent English communication skills.

Benefits

  • Remote work with a globally distributed team.
  • Opportunity to work on advanced multimodal AI and model-compression research in a fintech and digital-asset company.
  • Opportunity to publish findings in leading AI and machine-learning conferences.

Tech Stack

Categories

AI Research
Tether Operations Limited

About Tether Operations Limited

201-500 employees
Contact me