Mindrift

Freelance Agent Evaluation Engineer

Mindrift
Apply
18 hours ago
Remote, Italy +2 moreSenior

Responsibilities

  • Build realistic simulated developer environments containing codebases, infrastructure, tickets, documentation, conversations, and development history.
  • Design challenging tasks from intermediate environment states and define clear, solvable success criteria for AI coding agents.
  • Write functional and integration tests that accept all valid solutions and reject incorrect solutions.
  • Review agent solutions, analyze failures, and iterate on tasks and tests based on QA feedback.
  • Create evaluation datasets for assessing how well AI agents handle real-world developer tasks.

Requirements

  • At least 5 years of software-development experience.
  • Experience with Python and FastAPI.
  • Experience with JavaScript or TypeScript and React.
  • Experience with Docker, Postgres, Kafka, and Redis.
  • Experience writing functional and integration tests.
  • A Master’s degree in a listed technical or related field, or a Bachelor’s degree with 5 years of experience in the field.
  • At least 3 years of professional experience in related QA-automation/testing or cybersecurity roles or domains.
  • English proficiency at B2 level or higher.
  • CV submitted in English with English proficiency indicated.

Benefits

  • Part-time, remote, freelance project-based work that can fit around primary professional or academic commitments.
  • Opportunity to work on advanced AI projects and build portfolio experience.
  • Opportunity to influence how future AI models understand and communicate in the field.
  • Participation is project-based rather than permanent employment.
  • Payment is per accepted task, with rates depending on qualification tier and task efficiency.
Mindrift

About Mindrift

1,001-5,000 employees

Mindrift builds an expert-sourcing platform that connects domain specialists to project-based work training and evaluating generative AI models, including supervised fine-tuning, RLHF, evaluation, and red-teaming. It is built and operated by Toloka, part of Nebius Group, and run from Amsterdam, Netherlands. Work is fully remote and freelance, serving global technology companies developing and improving large AI systems.

Contact me