Mindrift

Freelance Agent Evaluation Engineer

Mindrift
Apply
10 hours ago
Remote, CanadaSenior

Responsibilities

  • Build realistic simulated developer environments containing codebases, infrastructure, tickets, documentation, conversations, and development history.
  • Design challenging coding tasks from intermediate environment states and define clear criteria for successful solutions.
  • Write functional and integration tests that accept valid agent solutions and reject incorrect ones.
  • Review agent solutions, analyze failures, and refine tasks and tests in response to QA feedback.
  • Evaluate and improve AI coding-agent performance through robust, fair task design.

Requirements

  • At least 5 years of software development experience.
  • Experience with Python, FastAPI, JavaScript, TypeScript, React, Docker, Postgres, Kafka, and Redis.
  • Experience writing functional and integration tests.
  • English proficiency at B2 level or higher.
  • A master's degree in Computer Science, Software Engineering, Data Science/Data Analytics, Artificial Intelligence/Machine Learning, Computational Linguistics/Natural Language Processing, Information Systems, or a related field.
  • A bachelor's degree is accepted with 5 years of experience in the field.
  • At least 3 years of professional experience in related roles or domains, specifically for QA-automation/testing or cybersecurity roles.

Benefits

  • Part-time, remote freelance work that can fit around primary professional or academic commitments.
  • Project-based participation rather than permanent employment.
  • Payment per accepted task, with an effective rate of up to $50/hr depending on qualification tier and task efficiency.
  • Opportunity to work on advanced AI projects, build portfolio experience, and influence future AI systems.
Mindrift

About Mindrift

1,001-5,000 employees

Mindrift builds an expert-sourcing platform that connects domain specialists to project-based work training and evaluating generative AI models, including supervised fine-tuning, RLHF, evaluation, and red-teaming. It is built and operated by Toloka, part of Nebius Group, and run from Amsterdam, Netherlands. Work is fully remote and freelance, serving global technology companies developing and improving large AI systems.

Contact me