arxivcs.CVcs.MM2026-07-07
WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation
Wei Dong, Tianyu Fu, Zhe Yu, Hanning Wang, Anyang Su, Zhizhou Fang, et al.
As web agents increasingly demonstrate capabilities in automated task execution, the development of robust evaluation frameworks for assessing their navigation and task completion performance has emerged as a critical research priority. However, existing benchmarks exhibit fundam…