Importance for Retention Is Not Importance for Acquisition: A Falsifiable Test in Neural Networks
Neural networks are empirically successful and mathematically only partially understood. This gap is often discussed as a single undifferentiated mystery, which makes the field feel either more solved or more mysterious than it actually is. We decompose the mystery into four distinct questions with different, unequal degrees of resolution: (1) why local, gradient-based optimization finds good solutions in a non-convex landscape; (2) whether a trained network's generalization can be certified rather than merely observed; (3) what a trained unit or circuit actually represents; and (4) whether these separate mathematical languages describe one underlying object. We survey the current state of each question, and identify a specific, underappreciated gap in the third: post-hoc, data-free measures of a unit's functional importance, such as the recently proposed HOPE framework, describe what a trained network now contains, but say nothing about when or how readily that content was acquired during training. We argue that importance for retention and importance for acquisition are conceptually distinct properties that existing post-hoc importance methods do not generally distinguish, and we propose a primary, falsifiable experiment that would test whether the two in fact coincide on networks whose target function is fully known, together with a second, cheaper, and independent experiment testing a related but distinct question about compression-based generalization certificates. Neither experiment requires new theory to run. We treat the first as the paper's central contribution and the second as a complementary check. A small, reduced-scale pilot of the primary experiment, run on a single-hidden-layer network rather than the full protocol proposed in the paper, is also reported; it shows a consistent, statistically significant association in every run, but at a scale too narrow to draw conclusions about the general claim, and we are explicit about why. We argue against premature claims of a single unifying mathematical object for learned knowledge until the full protocol, not this pilot, has actually been run.