Mindrift

Freelance Agent Evaluation Engineer

Mindrift
Apply
1 day ago
Remote, PhilippinesSenior

Responsibilities

  • Build realistic simulated developer environments with codebases, infrastructure, tickets, documentation, conversations, and development history.
  • Design challenging coding tasks from intermediate environment states and define clear, solvable success criteria.
  • Write functional and integration tests that accept all valid solutions and reject incorrect ones.
  • Review agent solutions, analyze failures, and iterate on tasks and tests based on QA feedback.
  • Create evaluation datasets for AI coding agents rather than writing most production code from scratch.

Requirements

  • At least 5 years of software development experience.
  • Experience writing functional and integration tests.
  • Proficiency with Python and FastAPI.
  • Proficiency with JavaScript or TypeScript and React.
  • Experience with Docker, Postgres, Kafka, and Redis.
  • Master’s degree in Computer Science, Software Engineering, Data Science/Data Analytics, Artificial Intelligence/Machine Learning, Computational Linguistics/Natural Language Processing, Information Systems, or a related field; a bachelor’s degree is accepted with 5 years of relevant experience.
  • At least 3 years of professional experience in related QA-automation/testing or cybersecurity roles or domains.
  • English proficiency at B2 level or higher.
  • Applicants must submit a CV in English and indicate their English proficiency.

Benefits

  • Part-time, remote freelance project that can fit around other professional or academic commitments.
  • Project-based participation rather than permanent employment.
  • Work on advanced AI projects and build portfolio experience.
  • Payment is per accepted task, with an effective rate of up to $30 per hour.
  • Opportunity to influence how future AI models understand and communicate in the candidate’s field.
Mindrift

About Mindrift

1,001-5,000 employees

Mindrift builds an expert-sourcing platform that connects domain specialists to project-based work training and evaluating generative AI models, including supervised fine-tuning, RLHF, evaluation, and red-teaming. It is built and operated by Toloka, part of Nebius Group, and run from Amsterdam, Netherlands. Work is fully remote and freelance, serving global technology companies developing and improving large AI systems.

Contact me