arxivcs.LGstat.ML2026-06-26
Benchmarking on Tasks That Matter: Dataset Selection for Preserving Model Rankings
Rostislav Gusev, Alexey Zaytsev
Benchmarks of machine learning models often include many datasets, making evaluation expensive. For efficiency, it is preferable to perform evaluations on small, representative datasets instead. The selection of such subsets typically relies on heuristics and is rarely analyzed f…