Tether Operations Limited

AI Research Engineer (Model Compression & Quantization)

Tether Operations Limited
Apply
4 months ago
Remote, WorldwideSenior

Responsibilities

  • Apply low-bit and mixed-precision quantization to LLMs, VLMs, and other multimodal generative AI models while preserving accuracy and output quality.
  • Use knowledge distillation to transfer capabilities from large teacher models to smaller student models across text, image, and audio inputs.
  • Implement pruning methods to remove redundant parameters and attention heads while maintaining task performance.
  • Build compression pipelines and establish metrics for model size, latency, throughput, memory use, accuracy, and fidelity.
  • Analyze efficiency-versus-accuracy trade-offs and propose improvements based on empirical results.
  • Research advanced methods such as adaptive pruning schedules and distillation with intermediate feature matching.
  • Identify and address production inference bottlenecks for low-memory, low-latency edge deployment.
  • Document methodologies, experiments, and results to support reproducibility and collaboration.
  • Author technical papers and publish model-compression research at leading AI conferences.

Requirements

  • Degree in Computer Science or a related field.
  • Strong AI R&D background with an excellent publication record, preferably including A* conference publications.
  • Hands-on experience with PyTorch or an equivalent deep-learning framework.
  • Hands-on experience with Quantization-Aware Training and Post-Training Quantization.
  • Research and hands-on experience with knowledge distillation for model compression.
  • Research and hands-on experience with model pruning for model compression.
  • Strong understanding of neural-network architectures and training, including transformers, LLMs, VLMs, backpropagation, optimization, and fine-tuning.
  • A PhD in NLP, Machine Learning, or a related field is preferred.
  • Familiarity with C++ is a plus, particularly for low-level quantization kernels or inference optimization.

Benefits

  • Remote work with a globally distributed team.
  • Opportunity to work on multimodal AI compression and efficient edge deployment in a digital-finance technology company.
  • Research and publication opportunities at leading conferences.

Tech Stack

Categories

AI ResearchML Engineering
Tether Operations Limited

About Tether Operations Limited

201-500 employees
Contact me