Mistral AI

Research Engineer, Inference Foundation

Mistral AI
Apply
8 hours ago
Paris, France +2 moreSenior

Responsibilities

  • Develop and fix the core inference engine and orchestrator, including feature selection, configuration, and performance tuning.
  • Own validated, regression-free serving-stack releases through automated performance gates and progressive rollout.
  • Contribute improvements and fixes upstream to open-source inference engines when appropriate.
  • Optimize serving efficiency across the fleet, including pod startup time, cold-cache regressions, caching, and offloading.
  • Optimize serving topology, including overlapping communication and transfers with computation and improving placement, connectivity, and routing.
  • Build serving infrastructure for reinforcement learning and post-training of frontier models.
  • Optimize inference performance across a broad range of workloads.

Requirements

  • Experience building and running ML or LLM services at scale with defined latency and availability targets.
  • Hands-on experience with inference engines such as vLLM, SGLang, TensorRT-LLM, or equivalent systems.
  • Strong understanding of prefill and decode, KV-cache behavior, batching, scheduling, speculative decoding, and parallelism strategies.
  • Familiarity with distributed and disaggregated serving architectures.
  • Ability to debug across CUDA/NCCL, kernels, containers, networking, and storage.
  • Proficiency with Python for systems tooling and backend services and experience with PyTorch.
  • Experience using Kubernetes to run infrastructure at scale.
  • Understanding of GPU and networking fundamentals, including CUDA runtime, NCCL, and InfiniBand/RDMA.
  • Preferred: vLLM or SGLang expertise, upstream contributions, demanding production experience, hardware-aware model optimization, MoE serving, CUDA/Triton kernel development, Nsight profiling, and Rust or C++ production experience.

Benefits

  • Benefits may include healthcare coverage, parental leave, retirement plans, relocation support, wellness programs, meal allowances, transportation allowances, and other country-specific perks.
Mistral AI

About Mistral AI

1,001-5,000 employees

Mistral AI builds foundation language models and full‑stack AI solutions for enterprises, offering APIs, on‑prem deployment, developer tools, and applications. The privately held company, founded in 2023 and headquartered in Paris, partners with organizations in finance, manufacturing, defense, healthcare, and the public sector to co‑create customized systems. It releases open‑source models alongside commercial offerings and makes its models available through major cloud marketplaces.

Contact me