
Research Engineer - Eval Platform
Mistral AI17 hours ago
Paris, FranceMid Level
Responsibilities
- Build systems that keep evaluation results reproducible and comparable over time.
- Run evaluations at scale across GPU clusters, from model serving through scoring.
- Develop APIs and dashboards that make evaluation results accessible, explorable, and trustworthy.
- Detect broken or noisy evaluations before they affect research decisions.
- Support evaluation of agentic, multi-turn, and tool-using models.
- Partner with researchers to turn evaluation needs into robust shared tooling.
Requirements
- Master's or PhD in Computer Science, or equivalent experience.
- At least 4 years of experience building production-grade software, ideally in large-scale ML codebases or distributed systems.
- Excellent Python skills and strong software-design instincts, including testing, code review, and CI/CD.
- Experience running workloads on GPU clusters such as Slurm, Kubernetes, or Ray.
- Familiarity with LLM inference and evaluation.
- Preferred qualifications include evaluation harness or benchmark experience, vLLM or SGLang experience, agentic or reinforcement-learning environment experience, experimentation statistics, and open-source ML tooling contributions.
Benefits
- Benefits may include healthcare coverage, parental leave, retirement plans, relocation support, wellness programs, meal allowances, transportation allowances, and other location-specific perks.
- Benefits vary by country.
Tech Stack
Categories
AI Research
About Mistral AI
Mistral AI builds foundation language models and full‑stack AI solutions for enterprises, offering APIs, on‑prem deployment, developer tools, and applications. The privately held company, founded in 2023 and headquartered in Paris, partners with organizations in finance, manufacturing, defense, healthcare, and the public sector to co‑create customized systems. It releases open‑source models alongside commercial offerings and makes its models available through major cloud marketplaces.