CORTEXA
← Browse
arxivcond-mat.dis-nncs.LGhep-lat2026-06-26

Spectral phase transitions and trainability in neural network learning dynamics

Chanju Park, Dario Bocchi, Francesco D'Amico, Biagio Lucini, Gert Aarts

The emergence of low-dimensional structures in the spectra of neural network weight matrices is a common empirical feature of trained models, but the dynamical origin of this phenomenon during learning remains an open problem. We formulate neural network training as the stochastic evolution of an initially random matrix ensemble, driven by stochastic gradient descent (SGD) updates that reshape the spectral bulk while amplifying signal strength. This induces a Baik-Ben Arous-Péché (BBP) transition during training, where isolated eigenvalues detach from the random bulk distribution, providing a dynamical framework for representation formation in high-dimensional learning dynamics. We demonstrate this in a solvable linear teacher-student model, where spectral evolution is analytically tractable and a phase diagram of trainability governed by the step size (or learning rate) and initial weight variance is obtained, and subsequently extend our formalism beyond the linear regime to nonlinear and stochastic settings. Numerical simulations in realistic settings support this picture, showing robust emergence of spectral alignment during training. Our results suggest that spectral analysis may provide a unified perspective of stochastic learning dynamics, linking trainability, optimisation hyperparameters, spectral phase transitions, and representation learning in neural networks.

View free PDFSource page

Related papers

arxivcs.LGcond-mat.dis-nnnlin.CDphysics.data-an2026-06-29

Scalar Representations of Neural Network Training Dynamics

Pedro Jiménez-González, Miguel C. Soriano, Lucas Lacasa

Training in artificial neural networks can be viewed as a trajectory evolving through a high-dimensional loss landscape. However, the large number of trainable parameters makes the direct analysis of these dynamics challenging. In this work, we treat such training trajectories as…

View free PDFSource page
arxivcs.LGcond-mat.dis-nncond-mat.stat-mech2026-07-11

Interpreting learning dynamics of autoencoders: Transient scaling and emerging concepts of the Ising model

Max Weinmann, Miriam Klopotek

We study how unsupervised autoencoders trained on microscopic spin configurations from the Ising model learn macroscopic, theory-relevant variables underlying the data-generating process. Without embedding domain knowledge, we mimic a typical discovery setting: We quantify learni…

View free PDFSource page
arxivcs.LGcond-mat.dis-nncs.AIstat.ML2026-06-26

How Width and Data Shape Generalization Scaling Laws in Quadratic Neural Networks

Julius Girardin, Emanuele Troiani, Yizhou Xu, Vittorio Erba, Florent Krzakala, Lenka Zdeborová

Understanding how performance scales jointly with model size and data is a central problem in modern machine learning. Existing theoretical works on scaling laws typically describe generalization as a function of data or compute, often in fixed-feature or infinite-width regimes a…

View free PDFSource page
arxivcs.LGcond-mat.dis-nn2026-07-09

An exact information theory of generalization phase transitions in Bayesian diffusion models

Henry Hunt, Mason Kamb, Surya Ganguli

How diffusion models circumvent the curse of dimensionality to learn complex distributions over high dimensional spaces from a finite training set, instead of memorizing it, remains a fundamental mystery. To address this, we introduce analytically tractable Bayesian information r…

View free PDFSource page
arxivcs.LGcond-mat.dis-nn2026-07-08

Explaining Near-Zero Hessian Eigenvalues Through Approximate Symmetries in Neural Networks

Marcel Kühn, Bernd Rosenow

The Hessian of the training loss governs the local geometry of the loss landscape, yet despite existing explanations for its largest eigenvalues, the origin of the vast multitude of vanishingly small eigenvalues remains elusive. We argue that the bulk consists of the weakly lifte…

View free PDFSource page
arxivcond-mat.dis-nncond-mat.stat-mechcs.LG2026-07-05

Broken Ergodicity and the Violation of the Fluctuation-Dissipation Theorem Lead to Generalization Beyond Overfitting in Machine Learning

Chan Li, Nigel Goldenfeld

The remarkable ability of modern neural networks to generalize improves with increasing network capacity, even when the number of model parameters or effective degrees of freedom exceeds the number of training data points. This phenomenon is all the more surprising given that gen…

View free PDFSource page