CORTEXA
← Browse
arxivcs.LGmath.ST2026-07-03

The Multiscale Single-Index Model: A Stylized Model for Hierarchical Feature Learning

Joan Bruna

We consider the Multiscale Single-Index Model (MSIM), first introduced in \cite{oymak2021learning}, as a stylized model for hierarchical learning with \emph{scale separation}. Each layer extracts a shared single-index feature at one physical scale and passes it to the next, thus defining a tractable setting in which to study how deep architectures learn multiscale representations. Under non-degeneracy and delocalization assumptions on the link function and planted features respectively, for fixed depth $K$ and local scale $d$, the first Wiener chaos of the target behaves as a perturbed spiked tensor, where the perturbation of order $d^{-1/2}$ comes from the non-linearity -- revealing the MSIM as a natural non-linear analogue of the Tensor PCA model \cite{montanari2014statistical}. While this perturbative picture is sufficient to enable efficient spectral recovery based on Tensor unfolding (as already observed in \cite{oymak2021learning}), it is not precise enough for the analysis of backpropagation gradient-based methods. In this work, we address this limitation by performing a fine-grained analysis of the Wiener chaos using Edgeworth expansions. In the first chaos, this gives a finite-rank hierarchy at scales $d^{-q/2}$. In higher chaoses, balanced flattenings exhibit staircase singular-value plateaus of size $d^{-ρ/2}$ and multiplicity $d^ρ$ under a natural higher-chaos non-cancellation condition. Using this higher-chaos structure, and under an additional slow Hermite-energy tail condition, we first establish shallow-network approximation lower bounds, quantifying the benefit of depth in this model. Next, and most importantly, we prove that online SGD on the correlation objective, where all layers evolve in the same timescale, achieves $1 - o_d(1)$ recovery with $n = \widetilde{O}( d^{K-1})$ samples, recovering the same sample complexity as in the linear counterpart.

View free PDFSource page

Related papers

arxivcs.LGmath.ST2026-06-26

Replica Symmetry Breaking and Algorithmic Thresholds in Empirical Risk Minimization under Multi-Index Model

Andrea Montanari, Kangjie Zhou

Modern machine learning models are trained by optimizing high-dimensional non-convex empirical risk functions. Such cost functions can have a multitude of local optima and yet, gradient-based optimization appears to converge to near-global optima. Within a simple supervised learn…

View free PDFSource page
arxivstat.MLcs.LGmath.ST2026-07-20

Mixing-Free and Signal-Optimal Learning of Gaussian Graphical Models from Glauber Dynamics

Vignesh Tirukkonda, Gautam Dasarathy

Gaussian graphical model selection is usually studied under independent sampling, but in many applications the data arise as a single trajectory of a dependent stochastic process. We study exact recovery of the graph from one trajectory of random-scan Gaussian Glauber dynamics. E…

View free PDFSource page
arxivmath.STcs.LGmath.PR2026-07-08

Any-Dimensional Learning by Sampling

Eitan Levin, Venkat Chandrasekaran

Many machine learning models are defined for inputs of different sizes, such as point clouds containing different numbers of points, sequences of tokens of different lengths, and graphs on different numbers of nodes. Such models are trained on finitely-many examples of necessaril…

View free PDFSource page
arxivcs.LGmath.OCmath.STstat.ML2026-07-02

Regularized Variational and Spectral Log-Density-Ratio Estimation in the Gaussian Location Model

Francis Bach

We study ridge-regularized log-density-ratio estimation in the Gaussian location model with a common covariance matrix. By affine invariance, the model is written as q $\sim$ N(0, I), p $\sim$ N($Δ$, I), with linear features, where $Δ$ is a mean vector. The variational estimator…

View free PDFSource page
arxivmath.STcs.LGstat.ML2026-06-30

On Optimal Data Splitting for Split Conformal Prediction

Sayan Das, Bahram Yaghooti, Todd A. Kuffner, Soumendra N. Lahiri

Conformal prediction and its variants, including the split conformal prediction, provide a distribution-free framework for uncertainty quantification by constructing prediction intervals or sets with finite-sample coverage guarantees. The statistical efficiency of these intervals…

View free PDFSource page
arxivstat.MLcs.AIcs.LGmath.STstat.CO2026-06-27

Perspectives on Latent Factor Indeterminacy and its Implications for Data Representation

Carel F. W. Peeters

The common factor analytic model is related to Helmholtz and Boltzmann machines, can be conceived as a linear autoencoder, or can be thought of as a single-hidden-layer generative neural network. We thus consider it a basal generative representation learner that can be used as a…

View free PDFSource page