Evaluation Rigor from Graph Neural Networks to Graph Foundation Models: A Systematic Review and a Four-Axis Reporting Standard
Sergei O. Kurashkin, Vadim S. Tynchenko, Aleksei S. Borodulin, Ahmad Hammoud, Connie Tee
Graph machine learning reports steady progress across node, graph, and link prediction, across temporal and hypergraph frontiers, and across the emerging class of graph foundation models. This review asks a prior question: when a method is reported to outperform the alternatives, how far does the evidence support the claim? We organize the answer around four axes of evaluation rigor: statistical rigor (seeds, dispersion, formal significance testing), baseline fairness (budget-parity tuning of trivial and structure-agnostic baselines), data integrity (leakage, duplication, negative sampling, contamination), and claim integrity (whether gains survive fair tuning and discriminative benchmarks). Drawing on a criterion-based corpus of 150 studies, of which 51 were read in full depth, we find a consistent picture. Only two of 25 methodologically central backbone studies apply a formal between-method significance test, and reported gains repeatedly shrink or disappear once a trivial baseline is tuned to parity, a leaked split is repaired, or a pretrained model is evaluated on unseen data. We argue that these failures share one cause: the saturation of benchmarks that can no longer discriminate between methods. The principal output is a minimum reporting standard, a concrete four-axis checklist that authors and reviewers can apply at negligible cost.