CORTEXA
← Browse
arxivstat.MLcs.LG2026-07-07

Width-Robust Learnability in Mean-Field Bayesian Neural Networks

Dmitry Vaintrob, Kaarel Hänni

Infinite-width limits are a standard way to reason about neural networks, but it is not automatic that the limiting learner has the same complexity-theoretic inductive bias as large finite networks. We study this question for Bayesian neural networks at the mean-field, or critical feature-learning, scaling. The central quantity is the \emph{reduced entropy} \[ s_\infty(y,\varepsilon)=\limsup_N -\frac{1}{N}\log π_N^0(L\le \varepsilon), \] the intensive prior cost of representing a target function $y$ to population mean-squared error $\varepsilon$. Our main result is a width-robust learnability theorem. At fixed depth, a family of Boolean-cube targets is learnable from polynomially many samples at infinite width if and only if it is learnable at polynomial width, if and only if its reduced entropy is polynomially bounded. Equivalently, up to polynomial slack in accuracy, the Bayesian mean-field learner generalizes exactly on the targets that can be represented by polynomial-size networks. The forward direction is proved by a form of subsampling: from the infinitely many hidden neurons in the mean-field solution, one can select polynomially many representatives and still preserve the learned function on every input simultaneously. At the critical scaling this subsampling has both an ``active'' component, which keeps the data-dependent low-dimensional statistics, and a ``lazy'' component, which resamples the entropy-dominated directions from the prior. Thus the infinite-width mean-field limit gives a clean analytic description of learning without introducing spurious width-dependent generalization power.

View free PDFSource page

Related papers

arxivcs.LGstat.ML2026-07-06

Minimum Block Width for Universal Approximation by Residual Neural Networks with Inner Width One

Qi Zhou, Xuan Zhou, Xiao-Song Yang

In this paper, we study the universal approximation property of residual neural networks, and obtain some new results. For input and output dimensions $d_x$ and $d_y$, and LeakyReLU, ReLU, ReLU-like activation functions, the upper and lower bounds of the minimum block width are e…

View free PDFSource page
arxivcs.LGstat.ML2026-07-08

LiST: Lipschitz Scaling Training for Robust and Calibrated Neural Networks

Arthur Chiron, Franck Mamalet, Thomas Massena, Thomas Deltort, Mathieu Serrurier

While accuracy, robustness, and calibration are all essential for reliable neural networks, they are often studied separately; developing models that satisfy all three simultaneously remains a central challenge. Lipschitz-constrained models guarantee robustness by design, yet the…

View free PDFSource page
arxivstat.MLcs.ITcs.LGmath.NA2026-07-12

Approximation of Analytic Functions by ReLU Neural Networks with Adjustable Depth and Width

Yanming Lai, Defeng Sun, Yang Wang

In contrast to most studies on neural network approximation theory that characterize results through a single parameter, such as the total number of network parameters, \cite{shen2020deep} pioneered the characterization of approximation rates as a joint function of the width para…

View free PDFSource page
arxivstat.MLcs.LG2026-07-02

Born Discrete, Made Smooth: Variational Formulation of Shallow Neural Networks

Matej Benko, Pierre Bousquet, Iwona Chlebicka, Błażej Miasojedow

Although neural networks are remarkably effective, their underlying optimization principles remain theoretically elusive, often characterized by non-convex landscapes and stochastic heuristics. In this work, we propose a paradigm shift by replacing the discrete training problem o…

View free PDFSource page
arxivcs.LGstat.ML2026-07-16

Probabilistic Physics-Informed Neural Networks for Estimating Heterogeneous Elastic Properties from Low-Resolution and Noisy Displacement Data

Tatthapong Srikitrungruang, Jaesung Lee

Estimating spatially heterogeneous elastic properties from low-resolution displacement measurements is a severely ill-posed inverse elasticity problem because low resolution obscures spatial details needed to distinguish heterogeneous property variations, and small measurement pe…

View free PDFSource page