CORTEXA
← Browse
arxivcs.LG2026-07-04

Validation-Induced Shapley Shifts: How Validation Structure Distorts Data Valuation

Yinan Shen, Ziao Yang, Hongfu Liu

Shapley values are widely used to attribute value to training data based on their marginal contribution to performance on a validation set. Existing practice often assumes these values are stable once the training data and model are fixed. In this work, we uncover a systematic vulnerability: even modest changes to the validation set, such as introducing noises, cause directional shifts in Shapley distributions. As noises are added, Shapley values of training samples compress toward zero. We trace this to a noise-induced neighborhood reshuffling effect: perturbations alter the local rank order between validation and training samples, flattening the valuation landscape. Using the KNN-Shapley framework, we show through synthetic and real data that these shifts are consistent and reproducible. Our findings challenge the assumption of Shapley stability and reveal a new axis of fragility in data valuation. We propose normalization and boundary-aware validation strategies to mitigate these distortions and enable more robust, interpretable valuation in machine learning marketplaces.

View free PDFSource page

Related papers

arxivstat.MLcs.ITcs.LGmath.CAmath.CO2026-07-01

Function-Counting Theory for Low-Dimensional Data Structures

Konstantin Häberle, Helmut Bölcskei

The success of deep learning models in classification and regression is widely attributed to the low-dimensional structure that real-world data tend to exhibit, despite their high-dimensional representation. This work attempts to provide a mathematical framework for binary classi…

View free PDFSource page
arxivcs.LGcs.AI2026-07-22

Synthetic minority data is redundant or invalid: a data-dependent validity theory and a de-biased test

Ahmad B. Hassanat, Ahmad S. Tarawneh, Ghada A. Altarawneh

For two decades, the standard remedy for class-imbalanced learning has been to fabricate synthetic minority examples, and the standard evidence of their validity has been a check that cannot fail: synthetic points are scored against the very data that generated them. We de-bias t…

View free PDFSource page
arxivphysics.chem-phcs.LGphysics.comp-phphysics.data-an2026-07-22

Hypothesis-and-Refinement Learning of Organic Structures from Multimodal Spectroscopic Data

Chengchun Liu, Zhiyuan Yan, Li Yuan, Hao Li, Boxuan Zhao, Yonghong Tian, et al.

Determining molecular structures from spectroscopic data remains fundamentally challenging because the inverse problem is intrinsically underdetermined: individual spectra are sparse, low-dimensional, and encode only partial structural evidence relative to the vast space of possi…

View free PDFSource page
arxivcs.LGcs.AI2026-07-02

Evolutionary Feature Engineering for Structured Data

Ege Onur Taga, Yilin Zhuang, M. Emrullah Ildiz, Petros Mol, Abhimanyu Das, Karthik Duraisamy, et al.

Large language models are increasingly used as open-ended search operators in evolutionary optimization. We introduce Evolutionary Feature Engineering (EFE), a framework for using LLM-based evolution to discover preprocessing transformations for structured data. EFE represents tr…

View free PDFSource page