Tether Operations Limited

AI Research Engineer (Model Compression & Quantization) - 100% Remote Worldwide

Tether Operations Limited
Apply
4 months ago
Remote, WorldwideSenior

Responsibilities

  • Design and deploy high-throughput, low-latency model-serving architectures with optimized memory usage across mobile, edge, and other resource-constrained environments.
  • Establish performance targets and monitor latency, throughput, memory consumption, error rates, and other production inference metrics.
  • Build and run controlled inference tests in simulated and live production environments, documenting results against established benchmarks.
  • Prepare test datasets and simulation scenarios for evaluating model performance on low-resource devices.
  • Analyze serving-pipeline efficiency and resolve bottlenecks involving batching, network delays, processing, and memory usage.
  • Integrate optimized serving and inference frameworks into production pipelines for edge and on-device applications.
  • Apply distributed inference techniques, including tensor parallelism, pipeline parallelism, and expert parallelism, to support large models on GPU clusters.
  • Iterate on model-serving and inference optimizations to improve real-world performance, reliability, scalability, and memory efficiency.

Requirements

  • Bachelor's degree in Computer Science or a related field; a PhD in NLP, Machine Learning, or a related field is preferred.
  • Solid AI R&D track record with strong publications in A* conferences preferred.
  • Knowledge of Metal Shading Language and ability to write custom compute shaders from scratch.
  • Proven experience with low-level kernel and inference optimization on mobile devices, including measurable improvements in latency, throughput, and memory footprint.
  • Deep understanding of modern model-serving architectures, inference engines, inference optimization, and efficient memory management.
  • Strong expertise writing GPU kernels for mobile devices such as smartphones.
  • Practical experience developing and deploying end-to-end inference pipelines on resource-constrained devices and edge platforms.
  • Ability to apply empirical research, design evaluation frameworks, and address latency, computational, and memory constraints.
  • Understanding of Diffusion Models and Vision Transformers, along with pruning, quantization, FlashAttention, KV caching, and speculative decoding such as Eagle.

Benefits

  • 100% remote work worldwide.
  • Opportunity to work on advanced AI serving and inference systems within a global fintech and digital-asset company.
  • Collaboration with cross-functional teams on production edge, on-device, and GPU-cluster applications.

Categories

Tether Operations Limited

About Tether Operations Limited

201-500 employees
Contact me