Mindrift

Freelance Agent Evaluation Engineer

Mindrift
Apply
18 hours ago
Remote, India +4 moreSenior

Responsibilities

  • Build realistic developer environments containing codebases, infrastructure, tickets, documentation, conversations, and development history.
  • Design tasks from intermediate environment states, including prompts and definitions of successful solutions.
  • Write functional and integration tests that accept valid agent solutions and reject incorrect ones.
  • Review agent solutions, analyze failures, and refine tasks and tests based on QA feedback.
  • Create fair and robust evaluations for AI coding agents without primarily writing the solutions from scratch.

Requirements

  • At least 5 years of software development experience.
  • At least 3 years of professional experience in related QA automation/testing or cybersecurity roles where applicable.
  • Experience with Python, FastAPI, JavaScript/TypeScript, React, Docker, Postgres, Kafka, and Redis.
  • Experience writing functional and integration tests.
  • Master's degree in computer science, software engineering, data science/data analytics, artificial intelligence/machine learning, computational linguistics/natural language processing, information systems, or a related field.
  • A Bachelor's degree is accepted with 5 years of experience in the field.
  • English proficiency at B2 level or higher.
  • Submit a CV in English and indicate English proficiency.

Benefits

  • Part-time, remote, freelance project-based work.
  • Flexible work that can fit around primary professional or academic commitments.
  • Work on advanced AI projects and gain portfolio experience.
  • Influence how future AI models understand and communicate in the field.
  • Paid per accepted task, with rates depending on qualification tier and efficiency, up to the equivalent of $30 per hour.
Mindrift

About Mindrift

1,001-5,000 employees

Mindrift builds an expert-sourcing platform that connects domain specialists to project-based work training and evaluating generative AI models, including supervised fine-tuning, RLHF, evaluation, and red-teaming. It is built and operated by Toloka, part of Nebius Group, and run from Amsterdam, Netherlands. Work is fully remote and freelance, serving global technology companies developing and improving large AI systems.

Contact me