CORTEXA
← Browse
arxivstat.MEstat.APstat.ML2026-07-09

Joint estimation of high-dimensional spiked covariance matrices via a partially shared subspace

Changwon Yoon, Minwoo Kim, Sungkyu Jung, Jeongyoun Ahn

Statistical analysis of high-dimensional data is often hampered by limited sample sizes, yet auxiliary datasets from related sources are often readily available. When two such datasets share part of their covariance structure, but not all of it, exploiting the shared part can substantially improve estimation. We propose a spiked covariance model that explicitly captures this partial sharing: two datasets share a subspace of unknown rank and arbitrary position in the spectrum, while each retains its own distinct spiked directions. The model treats the two datasets symmetrically and strictly generalizes existing models for shared covariance structure. We develop a complete estimation procedure that includes joint estimation of the shared subspace and its rank, a closed-form pooling weight for combining the two datasets, and asymptotic guarantees derived from random matrix theory in the proportional-growth regime. The framework also resolves a gap in contrastive dimension reduction by providing a principled estimator for high-dimensional settings. We illustrate the methodology on portfolio construction during the early COVID-19 pandemic and on contrastive analysis of brain tumor gene expression.

View free PDFSource page

Related papers

arxivstat.MLcs.LGmath.PRstat.APstat.COstat.ME2026-07-21

A Bayesian Framework for Built-in Input Dimension Reduction for Gaussian Process Modeling

Eric Herrison Gyamfi, Emily L. Kang, Bledar A. Konomi, Guang Lin

Gaussian process (GP) modeling is widely used in computational science and engineering. However, fitting a GP to high-dimensional inputs remains challenging due to the curse of dimensionality. While various methods have been proposed to reduce input dimensionality, they typically…

View free PDFSource page
arxivstat.MEcs.LGstat.APstat.COstat.ML2026-07-23

Distributional Determinantal Point Process for Repulsive Clustering of Distributions

Khai Nguyen, Yang Ni, Elizabeth Juarez-Colunga, Peter Mueller

We introduce the distributional determinantal point process (dDPP) as a novel repulsive point process whose atoms are probability distributions rather than points in a real space. The dDPP is constructed via an L-ensemble with a sliced Wasserstein (SW) kernel between distribution…

View free PDFSource page
arxivstat.APstat.MEstat.ML2026-07-07

Stochastic generator of trajectories from record data: application to the fluctuations of a glacier's frontal position from a sample of moraines

Megret Maud, Mike Pereira, Nicolas Eckert, Naveau Philippe, Jomelli Vincent

The record values theory study elements of a time series that exceed all previous observations, which are of particular interest in fields such as sports or climate science. In this paper, we propose a statistical method based on the construction of a Brownian stochastic simulato…

View free PDFSource page