Anyone AI

Senior Software Engineer – Open Source & SWE-Bench Evaluation

Anyone AI
Apply
8 hours ago
Remote, Spain +8 moreSenior

Responsibilities

  • Review ML challenges and determine whether they are well designed, technically solvable, and appropriately difficult.
  • Evaluate datasets for meaningful learnable signals, data quality issues, synthetic-data artifacts, leakage, contamination, and unintended shortcuts.
  • Review experiment designs, model-selection methods, evaluation metrics, improvement thresholds, and statistical significance.
  • Verify reproducibility across complete data, model, and evaluation pipelines, including CPU and GPU environments.
  • Provide clear recommendations to improve, recalibrate, or exclude problematic tasks.

Requirements

  • At least 3 years of hands-on applied machine learning experience.
  • Strong experience with ML experiment design, model selection, hyperparameter tuning, model evaluation, data preprocessing, and validation.
  • Strong understanding of train, validation, and test splits and the ability to identify data leakage, label noise, distribution shift, spurious correlations, feature leakage, and data contamination.
  • Experience determining whether performance improvements are statistically meaningful rather than random fluctuations, with strong knowledge of appropriate ML evaluation metrics.
  • Ability to debug ML workloads across CPU and GPU environments and provide clear written technical feedback.
  • Preferred experience includes Kaggle, DrivenData, or similar ML competitions; benchmark dataset or challenge design; data-centric AI; synthetic data generation and validation; statistical testing, confidence intervals, and effect sizes; ML evaluation pipelines; RLHF; AI model evaluation; or ML curricula and technical assessments.

Benefits

  • Remote work arrangement.
  • Part-time, project-based consulting engagement.

Categories

Contact me