Mindrift

Freelance Agent Evaluation Engineer

Mindrift
Apply
1 day ago
Remote, Turkey +5 moreSenior

Responsibilities

  • Build realistic simulated developer environments containing codebases, infrastructure, tickets, documentation, conversations, and development history.
  • Design challenging AI-agent tasks from intermediate environment states, including prompts and solvability criteria.
  • Write functional and integration tests that accept valid agent solutions and reject incorrect ones.
  • Review agent solutions, analyze failures, and refine tasks and tests based on QA feedback.
  • Create fair and robust evaluations that distinguish between strong and weak AI coding-agent performance.

Requirements

  • At least 5 years of software development experience.
  • At least 3 years of professional experience in related QA-automation/testing or cybersecurity roles where applicable.
  • Experience writing functional and integration tests.
  • Working knowledge of Python with FastAPI, JavaScript/TypeScript with React, Docker, Postgres, Kafka, and Redis.
  • Master’s degree in computer science, software engineering, data science/data analytics, artificial intelligence/machine learning, computational linguistics/NLP, information systems, or a related field; a bachelor’s degree is accepted with 5 years of field experience.
  • English proficiency at B2 level or higher.
  • Ability to evaluate complex coding-agent behavior and create tests that accommodate multiple valid solutions.

Benefits

  • Remote, part-time freelance project-based work.
  • Flexible participation designed to fit around primary professional or academic commitments.
  • Exposure to advanced AI projects and portfolio-building experience.
  • Opportunity to influence how future AI models understand and communicate in the field.
  • Paid per accepted task, with an effective rate of up to $50 per hour depending on qualification tier and productivity.
Mindrift

About Mindrift

1,001-5,000 employees

Mindrift builds an expert-sourcing platform that connects domain specialists to project-based work training and evaluating generative AI models, including supervised fine-tuning, RLHF, evaluation, and red-teaming. It is built and operated by Toloka, part of Nebius Group, and run from Amsterdam, Netherlands. Work is fully remote and freelance, serving global technology companies developing and improving large AI systems.

Contact me