CORTEXA
← Browse
arxivmath.STstat.MEstat.ML2026-07-03

Tightening Control in Neyman--Pearson Linear Classification

Yijian Huang

Neyman--Pearson classification prioritizes one class by constraining its accuracy above a prespecified level, and then takes the accuracy of the other class as the utility objective. This paradigm is well suited for disease screening and diagnosis, among other applications. Statistical learning under this framework is complicated since classifier performance determines its acceptability. Furthermore, no learned classifier that is consistent for the oracle classifier can guarantee satisfaction of the control constraint in finite samples. Classical learning theory targets a control-relaxed empirical utility maximization (EUM) classifier. However, even the EUM classifier fails to achieve the desired control level on average. We conjecture that this under-control phenomenon is a manifestation of the over-optimism bias well known in standard statistical learning, and develop asymptotic theory to confirm it. Motivated by this insight, we propose refined learning procedures under two accuracy control strategies for the prioritized class: one controlling accuracy in expectation and the other with high probability. We further develop training-data-based methods to predict and infer class-specific accuracies of the resulting classifiers. Simulation studies demonstrate favorable finite-sample performance, and we illustrate the proposed methods with an application to cancer detection.

View free PDFSource page

Related papers

arxivmath.STstat.MEstat.ML2026-07-20

How Fast Do Signatures Learn? Statistical Theory and Applications for Path Regression

Blanka Horvath, Wen Su, Wu Su, Binnan Wang, Ruixun Zhang

Many prediction and decision-making problems in operations research involve path-valued covariates -- data that evolve over time -- for which path signatures have become a canonical feature representation. Their use is justified by a universal approximation theorem, but this is a…

View free PDFSource page
arxivecon.EMcs.LGmath.STstat.MEstat.ML2026-07-20

Vector Search As Nearest Neighbor Matching: RAG-based Policy Learning in Causal Inference

Masahiro Kato, Taka Kato

We propose one-step and two-step methods for policy learning with retrieval-augmented generation (RAG). We formulate RAG-based action selection under the potential outcome framework. In the two-step method, vector search retrieves action-specific neighboring evidence in an embedd…

View free PDFSource page
arxivstat.MEcs.LGmath.STstat.ML2026-07-17

Aggregation of Statistical Evidence under Exchangeability

Antonin Schrab, Rajen Shah, Arthur Gretton, Ilmun Kim

We study aggregation of statistical evidence under unknown and potentially complex dependence using group-invariance. Building on permutation-based constructions that treat transformed datasets as exchangeable units, we aggregate evidence across statistics for each transformed da…

View free PDFSource page
arxivmath.STcs.LGstat.MEstat.ML2026-07-20

Unveiling Invariant and Transferable Latent Factors Across Heterogeneous Environments via ATLAS

Yihong Gu, Katherine Liao, Tianxi Cai

This paper considers a multi-environment factor model in which high-dimensional covariates are collected from heterogeneous environments, with auxiliary labels available in a subset of these environments. The joint distribution of the covariates may vary across environments, wher…

View free PDFSource page