20 hours ago
Remote, Argentina +2 moreSenior
Responsibilities
- Build realistic simulated developer environments containing codebases, infrastructure, tickets, documentation, conversations, and development history.
- Design challenging tasks from intermediate environment states, including prompts and criteria for successful solutions.
- Write functional and integration tests that accept valid solutions and reject incorrect ones.
- Review agent solutions, analyze failures, and iterate on tasks and tests based on QA feedback.
- Create evaluations that reveal meaningful differences between strong and weak AI coding-agent performance.
Requirements
- At least 5 years of software development experience.
- Experience with Python and FastAPI, JavaScript/TypeScript and React, Docker, Postgres, Kafka, and Redis.
- Experience writing functional and integration tests.
- Master’s degree in Computer Science, Software Engineering, Data Science/Data Analytics, Artificial Intelligence/Machine Learning, Computational Linguistics/NLP, Information Systems, or a related field; a bachelor’s degree is accepted with 5 years of relevant experience.
- At least 3 years of professional experience in related QA automation/testing or cybersecurity roles, as specified in the academic and professional experience requirements.
- English proficiency at B2 level or higher.
- Applicants must submit a CV in English and indicate their English proficiency.
Benefits
- Part-time, remote, freelance project-based work that can fit around other professional or academic commitments.
- Paid per accepted task, with compensation up to the equivalent of $30/hr depending on qualification tier and efficiency.
- Opportunity to work on advanced AI projects, build portfolio experience, and influence how future AI models perform in the field.
- Participation is project-based and not permanent employment.
Tech Stack
Categories
About Mindrift
Mindrift builds an expert-sourcing platform that connects domain specialists to project-based work training and evaluating generative AI models, including supervised fine-tuning, RLHF, evaluation, and red-teaming. It is built and operated by Toloka, part of Nebius Group, and run from Amsterdam, Netherlands. Work is fully remote and freelance, serving global technology companies developing and improving large AI systems.
