1 day ago
Remote, Turkey +5 moreSenior
Responsibilities
- Build realistic simulated developer environments containing codebases, infrastructure, tickets, documentation, conversations, and development history.
- Design challenging AI-agent tasks from intermediate environment states, including prompts and solvability criteria.
- Write functional and integration tests that accept valid agent solutions and reject incorrect ones.
- Review agent solutions, analyze failures, and refine tasks and tests based on QA feedback.
- Create fair and robust evaluations that distinguish between strong and weak AI coding-agent performance.
Requirements
- At least 5 years of software development experience.
- At least 3 years of professional experience in related QA-automation/testing or cybersecurity roles where applicable.
- Experience writing functional and integration tests.
- Working knowledge of Python with FastAPI, JavaScript/TypeScript with React, Docker, Postgres, Kafka, and Redis.
- Master’s degree in computer science, software engineering, data science/data analytics, artificial intelligence/machine learning, computational linguistics/NLP, information systems, or a related field; a bachelor’s degree is accepted with 5 years of field experience.
- English proficiency at B2 level or higher.
- Ability to evaluate complex coding-agent behavior and create tests that accommodate multiple valid solutions.
Benefits
- Remote, part-time freelance project-based work.
- Flexible participation designed to fit around primary professional or academic commitments.
- Exposure to advanced AI projects and portfolio-building experience.
- Opportunity to influence how future AI models understand and communicate in the field.
- Paid per accepted task, with an effective rate of up to $50 per hour depending on qualification tier and productivity.
Tech Stack
Categories
About Mindrift
Mindrift builds an expert-sourcing platform that connects domain specialists to project-based work training and evaluating generative AI models, including supervised fine-tuning, RLHF, evaluation, and red-teaming. It is built and operated by Toloka, part of Nebius Group, and run from Amsterdam, Netherlands. Work is fully remote and freelance, serving global technology companies developing and improving large AI systems.
