Tether Operations Limited

AI Inference Engineer QVAC (100% remote Worldwide)

Tether Operations Limited
Apply
6 months ago
Remote, WorldwideSenior

Responsibilities

  • Develop and optimize the C++ inference runtime powering QVAC’s local AI stack.
  • Deploy machine learning models to edge devices using llama.cpp and ggml.
  • Improve model startup time, memory efficiency, throughput, latency, runtime stability, and long-session reliability across hardware.
  • Port and enhance inference engines for different GPU architectures and integrate GPU frameworks.
  • Collaborate with researchers on coding, training, and transitioning models from research into production.
  • Integrate AI features into existing products and define core abstractions for future inference capabilities.

Requirements

  • Strong programming skills in C++.
  • Strong experience with llama.cpp and ggml inference engines.
  • Experience with at least one of CUDA, Vulkan, Metal, or OpenCL.
  • Good understanding of deep learning concepts and model architectures.
  • Experience with transformers, LLMs, and diffusion models.
  • Ability to rapidly learn new technologies and techniques.
  • A degree in Computer Science, AI, Machine Learning, or a related field, plus a solid track record in AI R&D.
  • Preferred experience training or fine-tuning LLMs, productionizing models, researching new model architectures, distributed systems, and JavaScript.

Benefits

  • 100% remote work worldwide.
  • Opportunity to build private, on-device AI infrastructure for QVAC and peer-to-peer AI products.
  • Work with a global, distributed team in the digital finance and AI technology space.

Tech Stack

Categories

Tether Operations Limited

About Tether Operations Limited

201-500 employees
Contact me