20 hours ago
Remote, France +2 moreSenior
Responsibilities
- Build realistic simulated developer environments containing codebases, infrastructure, tickets, documentation, and conversation history.
- Design coding tasks from intermediate environment states and define clear, solvable success criteria for AI agents.
- Write functional and integration tests that accept valid approaches and reject incorrect agent solutions.
- Review agent solutions, analyze failures, and iterate on tasks and tests using QA feedback.
- Create evaluation scenarios that reveal meaningful differences between strong and weak AI coding agents.
Requirements
- At least 5 years of software development experience.
- At least 3 years of professional experience in related roles or domains, specifically for QA automation/testing or cybersecurity roles.
- Master’s degree in Computer Science, Software Engineering, Data Science/Data Analytics, Artificial Intelligence/Machine Learning, Computational Linguistics/Natural Language Processing, Information Systems, or a related field; a bachelor’s degree is accepted with 5 years of field experience.
- Experience with Python, FastAPI, JavaScript/TypeScript, React, Docker, Postgres, Kafka, and Redis.
- Experience writing functional and integration tests.
- English proficiency at B2 level or higher.
- Submit a CV in English and indicate English proficiency.
Benefits
- Remote, part-time, freelance, project-based work that can fit around other professional or academic commitments.
- Paid per accepted task, with compensation up to the equivalent of $50/hr depending on qualification tier and task efficiency.
- Opportunity to work on advanced AI projects, build portfolio experience, and influence how future AI models understand and communicate in software development.
- Participation is project-based rather than permanent employment.
Tech Stack
Categories
About Mindrift
Mindrift builds an expert-sourcing platform that connects domain specialists to project-based work training and evaluating generative AI models, including supervised fine-tuning, RLHF, evaluation, and red-teaming. It is built and operated by Toloka, part of Nebius Group, and run from Amsterdam, Netherlands. Work is fully remote and freelance, serving global technology companies developing and improving large AI systems.
