CORTEXA
← Browse
arxivstat.MLcs.LG2026-07-21

Algebraic Signatures for Structural Learning in Probability Tensors

Akihiro Maeda, Shohei Hidaka, Satoshi Aoki

Algebraic statistics characterizes statistical models through polynomial constraints, but it has mainly been used for analytically specified model classes. This paper studies the inverse problem: identifying probabilistic structure from vanishing binomials observed in empirical probability tensors. We treat the vanishing binomials of a toric model as its algebraic signature, and turn the ideal-variety correspondence of algebraic statistics into an operational procedure for structural learning that identifies a model by signature matching without parameter estimation. By restricting attention to a computationally tractable class of configuration matrices, which we call {\it the Kronecker-stack class}, we make these signatures explicitly enumerable. Within this class we define minimum invariant constraint (MIC) as the atomic unit characterizing each signature and generalizing the notion of independence. We tested this approach employing MICs on synthetic data as well as on corpus-scale real language data. The results suggested the utility of the method, revealing that the identified rank-one structures correspond to interpretable sets of words. These results open up a new avenue for applying algebraic statistics to computational linguistics.

View free PDFSource page

Related papers

arxivcs.LGstat.ML2026-07-01

From Structural Equation Modelling to Double Machine Learning: Robustness Analysis for Survey-Based Research

Ka Ching Chan, Qiana Liu, Sanjib Tiwari, Ranga Chimhundu

Structural equation modelling (SEM) is widely used in survey-based business and information systems research to assess latent constructs and theory-driven structural relationships. However, SEM path significance is obtained within a particular model specification and may not show…

View free PDFSource page
arxivstat.MLcs.LG2026-07-07

Tensor Train Diffusion: Leveraging Low-Rank Structures for High-Dimensional Score-Based Sampling

Robert Gruhlke, Julius Berner, David Sommer, Lorenz Richter

Diffusion models offer a powerful framework for sampling from complex probability densities by learning to reverse a noising process. A common approach involves solving for the time-reversed stochastic differential equation (SDE), which requires the score function of the evolving…

View free PDFSource page
arxivstat.MLcs.LG2026-07-08

Tensorized algorithms and scalable filtering methods for hidden Markov and factorial hidden Markov models

Roxana Barrios, Ioannis Sgouralis

A common method for the representation and analysis of time-series data is the hidden Markov model (HMM), where each observation is associated with a hidden state that evolves over time. However, many real-world systems are influenced by multiple independent factors, which are mo…

View free PDFSource page
arxivstat.MLcs.ITcs.LGmath.CAmath.CO2026-07-01

Function-Counting Theory for Low-Dimensional Data Structures

Konstantin Häberle, Helmut Bölcskei

The success of deep learning models in classification and regression is widely attributed to the low-dimensional structure that real-world data tend to exhibit, despite their high-dimensional representation. This work attempts to provide a mathematical framework for binary classi…

View free PDFSource page