CORTEXA
← Browse
arxivcs.LGcond-mat.dis-nn2026-07-09

An exact information theory of generalization phase transitions in Bayesian diffusion models

Henry Hunt, Mason Kamb, Surya Ganguli

How diffusion models circumvent the curse of dimensionality to learn complex distributions over high dimensional spaces from a finite training set, instead of memorizing it, remains a fundamental mystery. To address this, we introduce analytically tractable Bayesian information restricted diffusion (BIRD) models, in which each pixel observes restricted information about noisy data. A BIRD model time-reverses diffusion by inferring which past training sample produced its current restricted observation using the Bayesian posterior. This model class generalizes existing analytical diffusion models that use spatially local information restriction. We show that spatially local BIRD models closely approximate trained diffusion models \textit{early in training}, across different architectures such as UNets and DiTs. Under minimal assumptions on the data distribution, we identify an information-theoretic phase boundary between memorization and generalization in the joint space of amount of training data, time in the reverse generative process, and amount of information restriction: a BIRD model memorizes when the mutual information between its restricted noisy observations and the training data exceeds the log number of training points, and it generalizes otherwise. Experiments across a range of datasets confirm our theoretically predicted location for the transition. We find that generation proceeds near the edge of memorization: both spatially local BIRD models and early-training diffusion models track the memorization-generalization phase boundary by increasingly restricting information over time. Overall, our results reveal a fundamental role for information restriction in generative AI to circumvent the curse of dimensionality.

View free PDFSource page

Related papers

arxivcond-mat.dis-nncs.LGhep-lat2026-06-26

Spectral phase transitions and trainability in neural network learning dynamics

Chanju Park, Dario Bocchi, Francesco D'Amico, Biagio Lucini, Gert Aarts

The emergence of low-dimensional structures in the spectra of neural network weight matrices is a common empirical feature of trained models, but the dynamical origin of this phenomenon during learning remains an open problem. We formulate neural network training as the stochasti…

View free PDFSource page
arxivquant-phcond-mat.dis-nncs.LG2026-06-29

Diffusion-warm sampling of the XY model enables fast thermalization at scale

Sehmimul Hoque, Roger Melko, Pooya Ronagh

We introduce a novel technique for scalable sampling of spin-system states with continuous symmetries using diffusion models. By applying our approach to the XY model, a fundamental continuous-spin model in condensed matter physics, we show that our technique addresses the shortf…

View free PDFSource page
arxivcond-mat.dis-nncond-mat.stat-mechcs.LG2026-07-05

Broken Ergodicity and the Violation of the Fluctuation-Dissipation Theorem Lead to Generalization Beyond Overfitting in Machine Learning

Chan Li, Nigel Goldenfeld

The remarkable ability of modern neural networks to generalize improves with increasing network capacity, even when the number of model parameters or effective degrees of freedom exceeds the number of training data points. This phenomenon is all the more surprising given that gen…

View free PDFSource page
arxivcs.LGcond-mat.dis-nncs.AIstat.ML2026-06-26

How Width and Data Shape Generalization Scaling Laws in Quadratic Neural Networks

Julius Girardin, Emanuele Troiani, Yizhou Xu, Vittorio Erba, Florent Krzakala, Lenka Zdeborová

Understanding how performance scales jointly with model size and data is a central problem in modern machine learning. Existing theoretical works on scaling laws typically describe generalization as a function of data or compute, often in fixed-feature or infinite-width regimes a…

View free PDFSource page
arxivcond-mat.dis-nncs.CLcs.LG2026-07-19

The Geometry of Semantic Space: A Continuous Geometric Framework for the Transformer Architecture

Zhihua Liang

We present a continuous geometric framework that models the discrete algebraic operations of the Transformer architecture as an integro-differential equation (IDE) on a semantic fiber bundle $\calE = \calM \times \R^d$. Beginning from a single geometric axiom -- that the token se…

View free PDFSource page
arxivcs.LGcond-mat.dis-nncond-mat.stat-mech2026-07-11

Interpreting learning dynamics of autoencoders: Transient scaling and emerging concepts of the Ising model

Max Weinmann, Miriam Klopotek

We study how unsupervised autoencoders trained on microscopic spin configurations from the Ising model learn macroscopic, theory-relevant variables underlying the data-generating process. Without embedding domain knowledge, we mimic a typical discovery setting: We quantify learni…

View free PDFSource page