
AI Inference Engineer QVAC (100% remote Worldwide)
Tether Operations Limited6 months ago
Remote, WorldwideSenior
Responsibilities
- Develop and optimize the C++ inference runtime powering QVAC’s local AI stack.
- Deploy machine learning models to edge devices using llama.cpp and ggml.
- Improve model startup time, memory efficiency, throughput, latency, runtime stability, and long-session reliability across hardware.
- Port and enhance inference engines for different GPU architectures and integrate GPU frameworks.
- Collaborate with researchers on coding, training, and transitioning models from research into production.
- Integrate AI features into existing products and define core abstractions for future inference capabilities.
Requirements
- Strong programming skills in C++.
- Strong experience with llama.cpp and ggml inference engines.
- Experience with at least one of CUDA, Vulkan, Metal, or OpenCL.
- Good understanding of deep learning concepts and model architectures.
- Experience with transformers, LLMs, and diffusion models.
- Ability to rapidly learn new technologies and techniques.
- A degree in Computer Science, AI, Machine Learning, or a related field, plus a solid track record in AI R&D.
- Preferred experience training or fine-tuning LLMs, productionizing models, researching new model architectures, distributed systems, and JavaScript.
Benefits
- 100% remote work worldwide.
- Opportunity to build private, on-device AI infrastructure for QVAC and peer-to-peer AI products.
- Work with a global, distributed team in the digital finance and AI technology space.