Mindrift

Freelance Agent Evaluation Engineer

Mindrift
Apply
1 day ago
Remote, Czechia +4 moreSenior

Responsibilities

  • Build realistic simulated developer environments with codebases, infrastructure, tickets, documentation, and conversations.
  • Design tasks from intermediate environment states, including prompts and objective solution criteria.
  • Write functional and integration tests that accept valid agent solutions and reject incorrect ones.
  • Review agent solutions, analyze failures, and iterate on tasks and tests based on QA feedback.
  • Create challenging evaluation datasets for AI coding agents rather than writing most production code from scratch.

Requirements

  • At least 5 years of software development experience.
  • A Master’s degree in a related field such as Computer Science, Software Engineering, Data Science, Artificial Intelligence, Computational Linguistics, or Information Systems; a Bachelor’s degree is accepted with 5 years of field experience.
  • At least 3 years of professional experience in related QA-automation/testing or cybersecurity roles or domains.
  • Experience writing functional and integration tests.
  • Experience with Python and FastAPI, JavaScript/TypeScript and React, Docker, Postgres, Kafka, and Redis.
  • English proficiency at B2 level or higher.
  • CV submitted in English with English proficiency indicated.

Benefits

  • Part-time, remote, freelance project-based work that can fit around primary professional or academic commitments.
  • Opportunity to work on advanced AI projects, enhance a portfolio, and influence how future AI models perform and communicate.
  • Paid per accepted task, with rates depending on qualification tier and task efficiency, up to the equivalent of $50 per hour.
  • Participation is project-based and not permanent employment.
Mindrift

About Mindrift

1,001-5,000 employees

Mindrift builds an expert-sourcing platform that connects domain specialists to project-based work training and evaluating generative AI models, including supervised fine-tuning, RLHF, evaluation, and red-teaming. It is built and operated by Toloka, part of Nebius Group, and run from Amsterdam, Netherlands. Work is fully remote and freelance, serving global technology companies developing and improving large AI systems.

Contact me