Mindrift

Freelance Agent Evaluation Engineer

Mindrift
Apply
18 hours ago
Remote, Japan +5 moreSenior

Responsibilities

  • Build realistic simulated developer environments with codebases, infrastructure, tickets, documentation, and conversations.
  • Design challenging developer tasks from intermediate environment states and define clear solution criteria.
  • Write functional and integration tests that accept valid agent solutions and reject incorrect ones.
  • Review agent solutions, analyze failures, and iterate on tasks and tests based on QA feedback.
  • Create fair and robust evaluations of AI coding agents rather than writing application code from scratch.

Requirements

  • At least 5 years of software development experience.
  • At least 3 years of professional experience in related QA automation/testing or cybersecurity roles, as specified for those domains.
  • Experience with Python and FastAPI, JavaScript/TypeScript and React, Docker, Postgres, Kafka, and Redis.
  • Experience writing functional and integration tests.
  • Master's degree in Computer Science, Software Engineering, Data Science/Data Analytics, Artificial Intelligence/Machine Learning, Computational Linguistics/NLP, Information Systems, or a related field; a bachelor's degree is accepted with 5 years of relevant experience.
  • English proficiency at B2 level or higher.
  • Candidates must submit a CV in English and indicate their English proficiency.

Benefits

  • Remote, part-time freelance project-based work that can fit around other professional or academic commitments.
  • Paid per accepted task, with rates based on qualification tier and task efficiency and up to the equivalent of $50 per hour.
  • Opportunity to work on advanced AI projects, build portfolio experience, and influence future AI model development.
  • Participation is project-based and not permanent employment.
Mindrift

About Mindrift

1,001-5,000 employees

Mindrift builds an expert-sourcing platform that connects domain specialists to project-based work training and evaluating generative AI models, including supervised fine-tuning, RLHF, evaluation, and red-teaming. It is built and operated by Toloka, part of Nebius Group, and run from Amsterdam, Netherlands. Work is fully remote and freelance, serving global technology companies developing and improving large AI systems.

Contact me