
LILT (Production)
Open Positions at LILT (Production)
10 open positions
Design and validate multilingual Terminal-Bench tasks that rigorously evaluate coding agents and large language models. This remote freelance role combines software engineering, benchmark development, deterministic verification, and Chinese Mandarin language expertise.
Design and validate Hindi-language Terminal-Bench tasks that rigorously evaluate coding agents and large language models on multilingual software challenges. This remote freelance role combines Python and shell-based implementation with deep expertise in Unicode, locale, and language-specific processing.
Design and validate multilingual Terminal-Bench tasks that rigorously evaluate coding agents and large language models across Serbian-language software workflows. This remote freelance role combines software engineering, benchmark development, language expertise, and deterministic verification.
Design and validate Japanese-language Terminal-Bench tasks that rigorously evaluate large language models on multilingual software challenges. This remote freelance role combines software engineering, benchmark development, multilingual text expertise, and human quality assurance.
Design and validate Korean-language Terminal-Bench tasks that evaluate coding agents and large language models on multilingual software challenges. This remote freelance role combines software engineering, benchmark construction, deterministic verification, and deep expertise in Korean language and locale-specific text handling.
Design and validate multilingual Terminal-Bench tasks that rigorously evaluate coding agents and large language models in Arabic and other language-sensitive software scenarios. This remote freelance role combines software engineering, benchmark creation, verification, and native-language expertise.
Design and validate multilingual Terminal-Bench tasks that rigorously evaluate coding agents and large language models in Turkish-language software environments. This remote freelance role combines software engineering, benchmark development, deterministic verification, and multilingual text expertise.
Design and validate multilingual Terminal-Bench tasks that rigorously evaluate coding agents and large language models across languages, locales, encodings, and terminal workflows. This remote freelance role is for a native Czech-speaking software engineer with strong Python and CLI experience.
Design and validate multilingual Terminal-Bench tasks that rigorously evaluate coding agents and large language models in German-language software environments. This remote freelance role combines software engineering, benchmark creation, deterministic verification, and native-language expertise.
Design and validate multilingual Terminal-Bench tasks that rigorously evaluate coding agents and large language models in Spanish-language environments. This remote freelance role combines software engineering, benchmark development, verifier scripting, and linguistic quality assurance.