Mindrift

Freelance Agent Evaluation Engineer

Mindrift
Apply
1 day ago
Remote, Japan +3 moreSenior

Responsibilities

  • Build realistic simulated developer environments containing codebases, infrastructure, tickets, documentation, conversations, and development history.
  • Create challenging coding-agent tasks from intermediate environment states, including prompts and solvability criteria.
  • Write functional and integration tests that accept valid solutions and reject incorrect ones.
  • Review agent solutions, analyze failures, and iterate on tasks and tests based on QA feedback.
  • Define fair and robust evaluation criteria for AI coding agents.

Requirements

  • At least 5 years of software development experience.
  • At least 3 years of professional experience in related QA automation/testing or cybersecurity roles.
  • Experience with Python and FastAPI.
  • Experience with JavaScript or TypeScript and React.
  • Experience with Docker, Postgres, Kafka, and Redis.
  • Experience writing functional and integration tests.
  • Master’s degree in computer science, software engineering, data science/data analytics, artificial intelligence/machine learning, computational linguistics/natural language processing, information systems, or a related field; a bachelor’s degree is accepted with 5 years of experience.
  • English proficiency at B2 level or higher.
  • CV submitted in English.

Benefits

  • Part-time, remote, freelance project-based work that can fit around other professional or academic commitments.
  • Work on advanced AI projects and build portfolio experience.
  • Opportunity to influence how future AI models understand and communicate in software development.
  • Paid per accepted task, with an effective rate of up to the equivalent of $50 per hour depending on qualification tier and efficiency.
  • Participation is project-based and not permanent employment.
Mindrift

About Mindrift

1,001-5,000 employees

Mindrift builds an expert-sourcing platform that connects domain specialists to project-based work training and evaluating generative AI models, including supervised fine-tuning, RLHF, evaluation, and red-teaming. It is built and operated by Toloka, part of Nebius Group, and run from Amsterdam, Netherlands. Work is fully remote and freelance, serving global technology companies developing and improving large AI systems.

Contact me