Mindrift

Freelance Agent Evaluation Engineer

Mindrift
Apply
20 hours ago
Remote, Poland +4 moreSenior

Responsibilities

  • Build realistic simulated developer environments containing codebases, infrastructure, tickets, documentation, conversations, and development history.
  • Design challenging tasks from intermediate environment states, including prompts, success criteria, and solvability requirements.
  • Write functional and integration tests that correctly accept valid AI-agent solutions and reject incorrect ones.
  • Review agent solutions, analyze failures, and iterate on tasks and tests using QA feedback.
  • Create evaluation datasets for assessing AI coding agents on real-world developer tasks.

Requirements

  • At least 5 years of software development experience.
  • Experience writing functional and integration tests.
  • Proficiency with Python and FastAPI, JavaScript or TypeScript and React, Docker, Postgres, Kafka, and Redis.
  • English proficiency at B2 level or higher.
  • A Master's degree in Computer Science, Software Engineering, Data Science or Data Analytics, Artificial Intelligence or Machine Learning, Computational Linguistics or NLP, Information Systems, or a related field.
  • A bachelor's degree is accepted with 5 years of experience in the field.
  • At least 3 years of professional experience in related roles or domains is specified for QA-automation/testing or cybersecurity roles.

Benefits

  • Remote, part-time freelance work that can fit around other professional or academic commitments.
  • Project-based participation with the opportunity to work on advanced AI projects and enhance a portfolio.
  • Opportunity to influence how future AI models understand and communicate in the field.
  • Payment is per accepted task, with the rate depending on qualification tier and task efficiency.
  • Participation is project-based rather than permanent employment.
Mindrift

About Mindrift

1,001-5,000 employees

Mindrift builds an expert-sourcing platform that connects domain specialists to project-based work training and evaluating generative AI models, including supervised fine-tuning, RLHF, evaluation, and red-teaming. It is built and operated by Toloka, part of Nebius Group, and run from Amsterdam, Netherlands. Work is fully remote and freelance, serving global technology companies developing and improving large AI systems.

Contact me