Mindrift

Freelance Agent Evaluation Engineer

Mindrift
Apply
1 day ago
Remote, India +5 moreSenior

Responsibilities

  • Build realistic virtual developer environments containing codebases, infrastructure, tickets, documentation, and conversation history.
  • Design tasks from intermediate environment states and define clear, solvable success criteria for AI coding agents.
  • Write functional and integration tests that accept valid solutions and reject incorrect ones.
  • Review agent solutions, analyze failures, and iterate on tasks and tests based on QA feedback.
  • Create fair and robust evaluations that reveal meaningful differences in AI coding-agent performance.

Requirements

  • At least five years of software development experience.
  • Experience writing functional and integration tests.
  • Familiarity with Python and FastAPI, JavaScript/TypeScript and React, Docker, Postgres, Kafka, and Redis.
  • English proficiency at B2 level or higher.
  • A master’s degree in Computer Science, Software Engineering, Data Science/Data Analytics, Artificial Intelligence/Machine Learning, Computational Linguistics/Natural Language Processing, Information Systems, or a related field.
  • A bachelor’s degree is accepted with five years of experience in the field.
  • At least three years of professional experience in related roles or domains, specifically including QA automation/testing or cybersecurity roles.

Benefits

  • Remote, part-time freelance work that can fit around other professional or academic commitments.
  • Project-based participation rather than permanent employment.
  • Paid per accepted task, with compensation up to the equivalent of $30 per hour depending on qualification tier and efficiency.
  • Opportunity to work on advanced AI projects, build portfolio experience, and influence future AI model development.
Mindrift

About Mindrift

1,001-5,000 employees

Mindrift builds an expert-sourcing platform that connects domain specialists to project-based work training and evaluating generative AI models, including supervised fine-tuning, RLHF, evaluation, and red-teaming. It is built and operated by Toloka, part of Nebius Group, and run from Amsterdam, Netherlands. Work is fully remote and freelance, serving global technology companies developing and improving large AI systems.

Contact me