CORTEXA
← Browse
arxivstat.MLmath.AP2026-06-26

Local Fokker--Planck Geometry for Score Estimation: Heat-Ball Mean-Value Representations and Exact High-Dimensional Sampling

Jiayao Bai, Lang Deng, Yi Du, Yifei Jia

Score-based generative models and Langevin samplers rely on estimating the score function $\nabla_x\log p_t(x)$ of a forward diffusion. Classically this is tractable when the drift is linear: the marginal density is Gaussian and the score is a global conditional expectation. For a general nonlinear, state-dependent drift the marginal density has no closed form, and existing methods--denoising score matching and global Fokker--Planck residual penalties--resort to global averaging that inflates estimation error in low-density regions precisely where accuracy is most critical. We address this by developing a local Fokker--Planck geometric framework that replaces global conditioning with local parabolic averaging. Our approach rests on three contributions. First, a time change to the cumulative-variance coordinate reduces the variable-coefficient Fokker--Planck equation to a standard inhomogeneous heat equation, on which we extend Evans' classical heat-ball monotonicity method to derive exact local mean-value representations for the score $\nabla_x\log p$ together with the density, log-density, and entropy density; local well-posedness is established under an explicit dimension-dependent drift budget. Second, for high-dimensional Monte Carlo evaluation of the resulting heat-ball integrals, we introduce the $κ$-measure and derive its exact factorized sampler with unit per-sample weight, $χ^2_2$ radial concentration. Third, the $r\to0$ limit of the heat-ball residual recovers the pointwise Fokker--Planck residual, showing that the local framework is a one-parameter generalization of global FP-residual methods, and that the DSM population minimizer is feasible for the heat-ball constraint at every scale. We validate the framework on 2D structured data on 256-dimensional MNIST, and on a dedicated sampler study confirming the concentration laws.

View free PDFSource page

Related papers

arxivstat.MEstat.APstat.ML2026-07-09

Joint estimation of high-dimensional spiked covariance matrices via a partially shared subspace

Changwon Yoon, Minwoo Kim, Sungkyu Jung, Jeongyoun Ahn

Statistical analysis of high-dimensional data is often hampered by limited sample sizes, yet auxiliary datasets from related sources are often readily available. When two such datasets share part of their covariance structure, but not all of it, exploiting the shared part can sub…

View free PDFSource page
arxivstat.MLcs.LGstat.COstat.ME2026-07-16

cGAP: Generalized Association Plots with HOMALS-Guided Heatmaps for Visualization of High-Dimensional Categorical Data

Chun-houh Chen, Shun-Chuan Chang, Chiun-How Kao, Yi-Ju Lee, Shang-Ying Shiu, Yin-Jing Tien, et al.

High-dimensional categorical data arise in genetics, biomedicine, and the social sciences, yet visualization tools for such data remain far less developed than those for continuous variables. Existing methods either scale poorly, rely heavily on low-dimensional displays detached…

View free PDFSource page
arxivmath.STmath.PRstat.ML2026-07-10

High-Dimensional Interpolators Can Be Fragile: Heavy Tails and High-Dimensional Large Deviations

Youheng Zhu, Yiping Lu

High-dimensional interpolation is common in modern machine learning, but its tail risk is less understood than its expected prediction risk. Existing theory shows that interpolating models can perform well in expectation, yet such guarantees do not determine the probability of ra…

View free PDFSource page