
Senior Inference Runtime Engineer
BitDeer Technologies Group2 months ago
Singapore, SingaporeSenior
Responsibilities
- Optimize prefill/decode scheduling, continuous batching, KV cache behavior, speculative decoding, long-context serving, and streaming smoothness.
- Tune and operate LLM inference runtimes for model-specific latency, throughput, GPU utilization, and cost efficiency.
- Profile bottlenecks across GPU memory, HBM bandwidth, NCCL/network, tokenizers, frontend/proxy layers, and model workers.
- Lead model onboarding, including runtime selection, tensor and pipeline parallelism, quantization, context length, and rollback strategy.
- Define runtime playbooks and safe defaults for reasoning, tool calling, multimodal serving, prompt caching, and provider-specific parameters.
- Partner with SRE and performance/evaluation engineers to turn benchmark findings into production runtime improvements.
Requirements
- 6+ years of systems, ML infrastructure, or high-performance backend engineering experience.
- Hands-on experience with LLM serving runtimes such as vLLM, Dynamo, SGLang, TensorRT-LLM, TGI, or Triton.
- Strong understanding of GPU memory, CUDA/NCCL basics, KV cache, batching, streaming, and distributed inference tradeoffs.
- Proficiency in Go or Python and ability to read runtime source code, profiling traces, and production metrics.
- Experience operating production inference services with strict latency, availability, and cost targets.
- Ability to translate low-level performance work into customer-visible reliability, latency, and margin improvements.
Benefits
- Inclusive and respectful workplace that values diverse backgrounds and perspectives.
- Opportunity to contribute directly to projects shaping Bitdeer’s digital asset and AI cloud businesses.
- Autonomy, personal accountability, fast growth, and learning opportunities.
- Training, mentoring, welfare benefits, networking opportunities, and involvement in new projects and process development.