CORTEXA
← Browse
arxivcs.LG2026-07-06

Measuring What Matters: A Unified Evaluation Framework for GNN Explainability

Francesco Paolo Nerini, Mirko Zaffaroni, Paolo Baracco, Gabriele Ciravegna, Alan Perotti

Graph eXplainable AI (G-XAI) is increasingly important for making Graph Neural Networks interpretable and accountable. While a growing number of explainers are available, choosing the right method and assessing the trustworthiness of its outputs remains unclear. Consistent evaluation practices and actionable guidance are still missing, hindering practical adoption. In this paper, we introduce a unified, quantitative benchmarking framework for G-XAI that requires no ground-truth assumptions. We formalize tabular explainability metrics for graph data, evaluating topological structure and node features as independent components. Our large-scale benchmarking study identifies explainers that consistently lie on the Pareto front across metric pairs and tasks, establishing robustly non-dominated solutions - while confirming that no single explainer achieves universal superiority. We distill our findings into actionable G-XAI usability guidelines to support Machine Learning practitioners in evaluating and deploying trustworthy GNN-based pipelines.

View free PDFSource page

Related papers

arxivcs.LGcs.AI2026-07-15Cited by 8

Towards a Unified Multidimensional Explainability Metric: Evaluating Trustworthiness in AI Models

Georgios Makridis, Georgios Fatouros, Athanasios Kiourtis, Dimitrios Kotios, Vasileios Koukos, Dimosthenis Kyriazis, et al.

In this paper, we present a comprehensive framework for assessing the explainability of various XAI methods, such as LIME and SHAP, across multiple datasets and machine learning models, with the ultimate goal of creating a unified multidimensional explainability score. Our method…

View free PDFSource page
arxivcs.LGcs.CV2026-07-23

Counterfactual Explainability Framework With CycleGAN And Counterfactual-Classifier Alignnment Score for Retinal Disease Classification

Kritanu Chattopadhyay, Sayanjit Singha Roy, Soumya Chatterjee

Automated detection of vision impairing retina-based ocular conditions from fundus images is important for early screening, timely referral and reducing dependency on specialist-only assessment, for which neural network-based deep learning (DL) models have been widely utilized. H…

View free PDFSource page
arxivcs.LG2026-06-30

TRIE: An Evaluation Framework for Stochastic PDE Surrogates

Bharat Srikishan, Javier E. Santos, Nikhil Muralidhar, Charles D. Young

Many scientific systems exhibit uncertainty from stochastic forcing, unresolved degrees of freedom, or imperfect observations, making reliable surrogate forecasting fundamentally distributional rather than pointwise. For such systems, deterministic neural surrogates fail to captu…

View free PDFSource page
arxivcs.LGcs.AI2026-07-23

GlucoTune: A Unified Framework for Blood Glucose Preprocessing, Forecasting, and Benchmarking in Diabetes

Davide Marelli, Giorgia Rigamonti, Mirko Paolo Barbato, Paolo Napoletano

Preprocessing blood glucose time-series data is a critical yet often overlooked step in developing data-driven methods for diabetes management, particularly for type 1 diabetes. The lack of standardized preprocessing workflows and evaluation protocols hinders reproducibility and…

View free PDFSource page