
AI Research Engineer (Model Compression & Quantization) - 100% Remote Worldwide
Tether Operations Limited4 months ago
Remote, WorldwideSenior
Responsibilities
- Design and deploy high-throughput, low-latency model-serving architectures with optimized memory usage across mobile, edge, and other resource-constrained environments.
- Establish performance targets and monitor latency, throughput, memory consumption, error rates, and other production inference metrics.
- Build and run controlled inference tests in simulated and live production environments, documenting results against established benchmarks.
- Prepare test datasets and simulation scenarios for evaluating model performance on low-resource devices.
- Analyze serving-pipeline efficiency and resolve bottlenecks involving batching, network delays, processing, and memory usage.
- Integrate optimized serving and inference frameworks into production pipelines for edge and on-device applications.
- Apply distributed inference techniques, including tensor parallelism, pipeline parallelism, and expert parallelism, to support large models on GPU clusters.
- Iterate on model-serving and inference optimizations to improve real-world performance, reliability, scalability, and memory efficiency.
Requirements
- Bachelor's degree in Computer Science or a related field; a PhD in NLP, Machine Learning, or a related field is preferred.
- Solid AI R&D track record with strong publications in A* conferences preferred.
- Knowledge of Metal Shading Language and ability to write custom compute shaders from scratch.
- Proven experience with low-level kernel and inference optimization on mobile devices, including measurable improvements in latency, throughput, and memory footprint.
- Deep understanding of modern model-serving architectures, inference engines, inference optimization, and efficient memory management.
- Strong expertise writing GPU kernels for mobile devices such as smartphones.
- Practical experience developing and deploying end-to-end inference pipelines on resource-constrained devices and edge platforms.
- Ability to apply empirical research, design evaluation frameworks, and address latency, computational, and memory constraints.
- Understanding of Diffusion Models and Vision Transformers, along with pruning, quantization, FlashAttention, KV caching, and speculative decoding such as Eagle.
Benefits
- 100% remote work worldwide.
- Opportunity to work on advanced AI serving and inference systems within a global fintech and digital-asset company.
- Collaboration with cross-functional teams on production edge, on-device, and GPU-cluster applications.