Mindrift

Freelance Agent Evaluation Engineer

Mindrift
Apply
20 hours ago
Remote, France +2 moreSenior

Responsibilities

  • Build realistic simulated developer environments containing codebases, infrastructure, tickets, documentation, and conversation history.
  • Design coding tasks from intermediate environment states and define clear, solvable success criteria for AI agents.
  • Write functional and integration tests that accept valid approaches and reject incorrect agent solutions.
  • Review agent solutions, analyze failures, and iterate on tasks and tests using QA feedback.
  • Create evaluation scenarios that reveal meaningful differences between strong and weak AI coding agents.

Requirements

  • At least 5 years of software development experience.
  • At least 3 years of professional experience in related roles or domains, specifically for QA automation/testing or cybersecurity roles.
  • Master’s degree in Computer Science, Software Engineering, Data Science/Data Analytics, Artificial Intelligence/Machine Learning, Computational Linguistics/Natural Language Processing, Information Systems, or a related field; a bachelor’s degree is accepted with 5 years of field experience.
  • Experience with Python, FastAPI, JavaScript/TypeScript, React, Docker, Postgres, Kafka, and Redis.
  • Experience writing functional and integration tests.
  • English proficiency at B2 level or higher.
  • Submit a CV in English and indicate English proficiency.

Benefits

  • Remote, part-time, freelance, project-based work that can fit around other professional or academic commitments.
  • Paid per accepted task, with compensation up to the equivalent of $50/hr depending on qualification tier and task efficiency.
  • Opportunity to work on advanced AI projects, build portfolio experience, and influence how future AI models understand and communicate in software development.
  • Participation is project-based rather than permanent employment.
Mindrift

About Mindrift

1,001-5,000 employees

Mindrift builds an expert-sourcing platform that connects domain specialists to project-based work training and evaluating generative AI models, including supervised fine-tuning, RLHF, evaluation, and red-teaming. It is built and operated by Toloka, part of Nebius Group, and run from Amsterdam, Netherlands. Work is fully remote and freelance, serving global technology companies developing and improving large AI systems.

Contact me