CORTEXA
← Browse
arxivstat.MLcs.LG2026-07-02

Born Discrete, Made Smooth: Variational Formulation of Shallow Neural Networks

Matej Benko, Pierre Bousquet, Iwona Chlebicka, Błażej Miasojedow

Although neural networks are remarkably effective, their underlying optimization principles remain theoretically elusive, often characterized by non-convex landscapes and stochastic heuristics. In this work, we propose a paradigm shift by replacing the discrete training problem of shallow neural networks with a well-posed continuum variational surrogate. We identify a family of $λ$-convex functionals over parameter densities in weighted Sobolev spaces and prove that these variational problems are globally well-posed, stable, and exhibit unexpected almost $C^3$ regularity. Unlike existing Wasserstein-based or Mean-Field approaches, which often face limited regularity and discretization challenges, our formulation provides direct access to elliptic regularity and convex analysis. This allows us to prove that the optimal parameter density can be obtained by solving a single linear system, bypassing iterative optimization entirely. We establish explicit generalization error controls at a rate of $1/α$ relative to the regularization parameter, and prove that finite-width networks of size $N$ achieve the continuum optimum at an $O(1/N)$ rate. This perspective bridges the gap between the Neural Tangent Kernel (NTK) and feature-learning regimes, providing a principled framework for understanding over-parameterization through the lens of variational calculus.

View free PDFSource page

Related papers

arxivcs.LGstat.ML2026-07-16

Probabilistic Physics-Informed Neural Networks for Estimating Heterogeneous Elastic Properties from Low-Resolution and Noisy Displacement Data

Tatthapong Srikitrungruang, Jaesung Lee

Estimating spatially heterogeneous elastic properties from low-resolution displacement measurements is a severely ill-posed inverse elasticity problem because low resolution obscures spatial details needed to distinguish heterogeneous property variations, and small measurement pe…

View free PDFSource page
arxivcs.LGstat.ML2026-07-06

Minimum Block Width for Universal Approximation by Residual Neural Networks with Inner Width One

Qi Zhou, Xuan Zhou, Xiao-Song Yang

In this paper, we study the universal approximation property of residual neural networks, and obtain some new results. For input and output dimensions $d_x$ and $d_y$, and LeakyReLU, ReLU, ReLU-like activation functions, the upper and lower bounds of the minimum block width are e…

View free PDFSource page
arxivstat.MLcs.ITcs.LGmath.NA2026-07-12

Approximation of Analytic Functions by ReLU Neural Networks with Adjustable Depth and Width

Yanming Lai, Defeng Sun, Yang Wang

In contrast to most studies on neural network approximation theory that characterize results through a single parameter, such as the total number of network parameters, \cite{shen2020deep} pioneered the characterization of approximation rates as a joint function of the width para…

View free PDFSource page
arxivcs.LGstat.ML2026-07-08

LiST: Lipschitz Scaling Training for Robust and Calibrated Neural Networks

Arthur Chiron, Franck Mamalet, Thomas Massena, Thomas Deltort, Mathieu Serrurier

While accuracy, robustness, and calibration are all essential for reliable neural networks, they are often studied separately; developing models that satisfy all three simultaneously remains a central challenge. Lipschitz-constrained models guarantee robustness by design, yet the…

View free PDFSource page