Tether Operations Limited

AI Research Engineer (Kernel & Inference Optimization) - 100% Remote Worldwide

Tether Operations Limited
Apply
18 hours ago
Remote, WorldwideSenior

Responsibilities

  • Design and deploy model-serving architectures optimized for throughput, latency, and memory usage across resource-constrained, edge, and on-device environments.
  • Build, execute, and monitor controlled inference tests in simulated and production environments using latency, throughput, memory, and error-rate metrics.
  • Create test datasets, simulation scenarios, benchmarks, and measurable evaluation criteria for real-world deployment conditions.
  • Analyze serving-pipeline efficiency and resolve batching, networking, memory, and computational bottlenecks.
  • Develop and optimize distributed inference engines using tensor, pipeline, and expert parallelism for large GPU clusters.
  • Collaborate with cross-functional teams to integrate optimized serving and inference frameworks into production systems.
  • Document results, define success metrics, and iteratively improve inference performance, scalability, reliability, and memory efficiency.

Requirements

  • Bachelor’s degree in Computer Science or a related field is required; a PhD in NLP, Machine Learning, or a related field is preferred.
  • Strong AI R&D track record, ideally including publications at A* conferences.
  • Knowledge of Metal Shading Language and the ability to write custom compute shaders from scratch.
  • Proven experience with low-level kernel optimization and inference optimization on mobile devices.
  • Demonstrated measurable improvements in inference latency, throughput, and memory footprint for domain-specific applications.
  • Deep understanding of model-serving architectures, inference optimization, memory management, and resource-constrained deployment.
  • Strong experience writing GPU kernels for smartphones and understanding model-serving frameworks and engines.
  • Practical experience developing and deploying end-to-end inference pipelines on constrained devices and edge platforms.
  • Ability to apply empirical research, design evaluation frameworks, diagnose bottlenecks, and iterate on optimization strategies.
  • Understanding of distributed inference systems, tensor parallelism, pipeline parallelism, expert parallelism, diffusion models, vision transformers, pruning, quantization, FlashAttention, KV cache, and speculative decoding such as Eagle.

Benefits

  • 100% remote work worldwide.
  • Opportunity to work on advanced AI systems and digital-finance technology with a globally distributed team.

Categories

AI ResearchML Engineering
Tether Operations Limited

About Tether Operations Limited

201-500 employees

Tether Operations Limited develops and issues reserve-backed digital tokens used by exchanges, wallets, payment processors, and merchants, most notably the USDT stablecoin available across multiple blockchains. Founded in 2014 and privately held, it earns revenue from token issuance, reserve management, and related services, and has expanded into tokenization, peer-to-peer communications (Keet), and Bitcoin mining. USDT is widely used in crypto trading and cross-border transfers, including in emerging markets.

Contact me