CORTEXA
← Browse
arxivmath.STecon.EMstat.MEstat.ML2026-07-06

Stabilized Higher-Order Influence Functions: Statistical Theory of a Class of Bilinear Forms

Na Liu, Chang Li, Yujia Gu, Lin Liu

Higher-order influence functions, introduced in a series of articles (Robins et al., 2008, 2009a; van der Vaart, 2014; Robins et al., 2016, 2023; Liu et al., 2017), are a unified framework for constructing rate-optimal point estimates of a class of statistical functionals under various complexity-reducing assumptions on the posited statistical model that generates the observed data. Although higher-order (influence functions) estimators are theoretically appealing, they have very limited practical uptake compared to their first-order counterparts. The original higher-order estimators proposed in Robins et al. (2008) and Robins et al. (2017) involve nonparametric density estimation of multi-dimensional covariates, a highly nontrivial statistical and computational problem on its own. The density estimator is, in turn, used in the evaluation of the inverse population Gram matrix $Ω$ of a set of $k$-dimensional basis transformations of covariates. There, $k$ is allowed to be as large as $o (n^2)$. To partially address this potential shortcoming, Liu et al. (2017) restrict $k$ to $o (n)$ and instead estimate $Ω$ directly using the inverse sample Gram matrix estimator, but computed from an independent sample often obtained by sample-splitting. Liu et al. (2017) refer to this alternative estimator as the empirical higher-order estimator. Although the empirical higher-order estimator bypasses density estimation, it suffers from numerical instability due to inverting a large-dimensional sample Gram matrix. In this article, for a class of bilinear forms/functionals that often appear in substantive fields, we propose a new stabilized higher-order estimator without sample splitting, which exhibits more stable finite-sample performance compared to the empirical higher-order estimator. We also prove that this new class of higher-order estimators enjoys similar statistical guarantees.

View free PDFSource page

Related papers

arxivecon.EMcs.LGmath.STstat.MEstat.ML2026-07-20

Vector Search As Nearest Neighbor Matching: RAG-based Policy Learning in Causal Inference

Masahiro Kato, Taka Kato

We propose one-step and two-step methods for policy learning with retrieval-augmented generation (RAG). We formulate RAG-based action selection under the potential outcome framework. In the two-step method, vector search retrieves action-specific neighboring evidence in an embedd…

View free PDFSource page
arxivecon.EMmath.STstat.MEstat.ML2026-07-07

Factor-Augmented Machine Learning Panel Regressions

Andrii Babii, Luca Barbaglia, Eric Ghysels, Jonas Striaukas

This paper develops the asymptotic theory for high-dimensional panel data regressions in settings with cross-sectionally dependent errors driven by common shocks. We consider a factor-augmented sparse-group LASSO estimator that combines MIDAS aggregation with latent factors. The…

View free PDFSource page
arxivmath.STstat.MEstat.ML2026-07-20

How Fast Do Signatures Learn? Statistical Theory and Applications for Path Regression

Blanka Horvath, Wen Su, Wu Su, Binnan Wang, Ruixun Zhang

Many prediction and decision-making problems in operations research involve path-valued covariates -- data that evolve over time -- for which path signatures have become a canonical feature representation. Their use is justified by a universal approximation theorem, but this is a…

View free PDFSource page
arxivstat.MEmath.STstat.COstat.ML2026-07-24

The V-fold jackknife for semiparametric inference: variance estimation, confidence intervals, and simultaneous confidence bands

Yi Li, Ashkan Ertefaie, Mark van der Laan

For decades, the bootstrap has been a default tool for statistical inference because of its broad applicability and minimal analytic requirements. Although its validity is well understood for smooth parametric estimators, its theoretical properties for many modern semiparametric…

View free PDFSource page
arxivstat.MEmath.STstat.ML2026-07-02

Cross-Audit Projection for Model Risk Prediction

Yijian Huang

For training-data-based model risk prediction, $K$-fold cross-validation~(CV) is widely used to mitigate the well-known over-optimism of the empirical risk and is often regarded as reliable. However, for binary classification via empirical risk minimization, our numerical studies…

View free PDFSource page