Mindrift

Freelance Agent Evaluation Engineer

Mindrift
Apply
20 hours ago
Remote, Italy +2 moreSenior

Responsibilities

  • Build realistic simulated developer environments containing codebases, infrastructure, tickets, documentation, and conversations.
  • Design tasks from intermediate environment states, including prompts, success criteria, and solvability checks.
  • Write functional and integration tests that verify AI agent solutions while accepting valid implementation approaches.
  • Review agent solutions, analyze failures, and refine tasks and tests based on QA feedback.
  • Create robust evaluations that distinguish between correct and incorrect AI coding-agent behavior.

Requirements

  • 5+ years of software development experience.
  • At least 3 years of professional experience in related roles or domains, specifically for QA automation/testing or cybersecurity roles.
  • Experience writing functional and integration tests.
  • Experience with Python and FastAPI, JavaScript/TypeScript and React, Docker, Postgres, Kafka, and Redis.
  • Master’s degree in Computer Science, Software Engineering, Data Science/Data Analytics, Artificial Intelligence/Machine Learning, Computational Linguistics/Natural Language Processing, Information Systems, or a related field; a bachelor’s degree is accepted with 5 years of experience.
  • English proficiency at B2 level or higher.

Benefits

  • Remote, part-time, freelance project-based work.
  • Flexible participation that can fit around primary professional or academic commitments.
  • Opportunity to work on advanced AI projects and enhance a professional portfolio.
  • Opportunity to influence how future AI models understand and communicate in the candidate’s area of expertise.
  • Paid per accepted task, with rates depending on qualification tier and completion efficiency; up to the equivalent of $40/hr.
Mindrift

About Mindrift

1,001-5,000 employees

Mindrift builds an expert-sourcing platform that connects domain specialists to project-based work training and evaluating generative AI models, including supervised fine-tuning, RLHF, evaluation, and red-teaming. It is built and operated by Toloka, part of Nebius Group, and run from Amsterdam, Netherlands. Work is fully remote and freelance, serving global technology companies developing and improving large AI systems.

Contact me