CORTEXA
← Browse
arxivcs.LGphysics.data-an2026-06-28

Anti-Collapse Dynamics and the Emergence of Multi-Time-Scale Learning in Recurrent Neural Networks

Lorenzo Livi

Long-range learning is hard for recurrent networks trained with stochastic gradient descent, because the influence of a past input fades with the lag $\ell$, and if it fades too fast the dependence cannot be learned from finite data. This fade is captured by an envelope $f(\ell)$. An exponential fade makes the data needed to learn a lag-$\ell$ dependence grow exponentially, putting long horizons out of reach; a power-law fade keeps the cost polynomial. We show that the asymptotic decay class of $f(\ell)$ is not fixed by the architecture. Instead, it emerges from the coupling between the state dynamics and parameter dynamics, settling into either a collapsed regime (fast, exponential forgetting) or an extended, anti-collapsed regime (slow, power-law forgetting). The intuition is a competition within these coupled dynamics. Training drives the network's effective time scales toward short ones, while rare, heavy-tailed fluctuations of the learning dynamics push a few of them to very long values. The extended regime survives only when these heavy-tailed pushes are strong enough to balance the pull. We make this mathematically precise with a coarse-grained stochastic process and prove exactly when the extended regime exists. A single exponent, the spectral exponent~$β$, then governs both the spread of time scales and how slowly the network forgets. Realizing the regime in practice needs one more ingredient: the joint action of the architecture and the optimizer must be able to hold such a broad spread. A network whose capacity to generate broad time-scale spectra is severely constrained still collapses, even when supplied with strong heavy-tailed forcing. Heavy-tailed fluctuations thus act not as noise to be suppressed, but as the mechanism that sustains long-range learning.

View free PDFSource page

Related papers

arxivcs.LGcond-mat.dis-nnnlin.CDphysics.data-an2026-06-29

Scalar Representations of Neural Network Training Dynamics

Pedro Jiménez-González, Miguel C. Soriano, Lucas Lacasa

Training in artificial neural networks can be viewed as a trajectory evolving through a high-dimensional loss landscape. However, the large number of trainable parameters makes the direct analysis of these dynamics challenging. In this work, we treat such training trajectories as…

View free PDFSource page
arxivphysics.soc-phcs.LGphysics.data-an2026-07-21

Tensor Network Machine Learning for Wildfire Susceptibility Mapping: from Grokking Dynamics to Quantum Mixedness of Class Representations

Domenico Pomarico, Alessandra Costantino, Gabriel Ramirez Sanchez, Loredana Bellantuono, Davide D' Alò, Mario Elia, et al.

A quantum-inspired tensor network framework for wildfire susceptibility classification in the Gargano region is introduced, leveraging AlphaEarth embeddings and Matrix Product State models. The approach combines scalable geospatial representations with an interpretable quantum ma…

View free PDFSource page
arxivcs.LGcond-mat.softcs.AIphysics.data-anstat.ML2026-07-21

Deep learning-based prediction of time-resolved adhesive forces in viscoelastic Hertzian contacts

Ali Maghami, Merten Stender, Michele Ciavarella, Antonio Papangelo

Fast prediction of the response of adhesive soft viscoelastic contacts represents a current challenge in soft robotics and for gripping and manipulation tasks. Determining the complete time-resolved force trajectory requires full numerical simulations, whose computational cost is…

View free PDFSource page
arxivcs.LGphysics.data-anphysics.flu-dyn2026-06-25

Kolmogorov Arnold networks (KAN) for aerodynamic prediction: a comparison with MLPs and GNNs

Miguel Jaraiz, Fermin Gutierrez, Pablo Yeste, Miguel Sánchez-Domínguez, Eusebio Valero, Gonzalo Rubio, et al.

Kolmogorov Arnold networks (KAN) have recently been introduced as a (deep) neural network architecture whose trainable parameters adapt the activation functions, instead of the coefficients of the affine transformations at the core of traditional architectures such as deep multil…

View free PDFSource page
arxivcs.LGmath.NAphysics.chem-phphysics.comp-phphysics.data-an2026-07-08

Higher-Order Geometric Updates for Levenberg-Marquardt Method via Riemann Normal Coordinates

Jianing Liu, Dong H. Zhang

Nonlinear least-squares optimization is central to regression, physics-informed neural networks, and other machine-learning tasks. Such problems have a natural geometric interpretation, model predictions form a manifold in data space, while the chosen parameterization can introdu…

View free PDFSource page
arxivq-bio.NCcs.LGphysics.data-anq-bio.QM2026-07-22

Computer Vision Based Neurology Brain Activity Rejection Architecture and Implementation

Zag ElSayed, Nathan Suer, Grace Westerkamp, Jack Yanchen Liu, Makoto Miyakoshi, Craig Erickson, et al.

The electroencephalogram (EEG) is a valuable and widely applied tool for investigating brain disorders and behavioral changes. It offers a minimally restrictive and non-invasive method. However, challenges in using EEG for cognitive development studies include temporal resolution…

View free PDFSource page