CORTEXA
← Browse
arxivcs.LGq-bio.GN2026-07-19

A multiverse-consensus pipeline for reproducible feature selection in untargeted LC-MS metabolomics

Mohammed Saeed Al-Huraibi, Ihsan Yozgat, Ahmet Kaplan

Background: Untargeted LC-MS metabolomics requires a long chain of preprocessing decisions, each with several equally defensible options. Analysts typically commit to one pipeline and report the resulting feature shortlist. How strongly that shortlist depends on choices that were never varied stays invisible. Results: We adapt multiverse analysis to untargeted metabolomics feature selection. We present an auditable, configuration-driven pipeline that (i) applies a ten-stage quality-control filter cascade in which every feature's fate is logged, and (ii) runs the downstream analysis as a multiverse over four contrasting preprocessing philosophies, each combined with four feature-ranking methods under bootstrap stability selection and label-permutation testing. Only features recurring across paths enter a tiered consensus. On a demonstration dataset of five breast-cancer cell lines (30,370 detected features), the four single pipelines individually returned shortlists of 4-20 features whose pairwise agreement was as low as Jaccard = 0.05. The multiverse consensus retained 15 features (>=2/4 paths), of which one recurred across all four, although two paths (sharing normalization and drift-correction methods) dominate the consensus. A pipeline-wide label-permutation test found no false discoveries in 50 null permutations. Conclusions: Reporting only preprocessing-robust features, with a complete kept/dropped audit trail, converts hidden analytical degrees of freedom into an explicit, inspectable output. We discuss scope and limitations, including single-batch design and the need for independent validation.

View free PDFSource page

Related papers

arxivq-bio.GNcs.AIcs.LG2026-07-21

Causal dictionary learning reveals and validates transcription-factor binding features in genomic language models

Sarwan Ali

Genomic language models achieve strong performance across regulatory-genomics tasks, yet what these models internally represent remains opaque, and the field lacks a principled procedure for verifying that an apparent ``concept'' inside a model is real rather than an artifact of…

View free PDFSource page
arxivcs.LGq-bio.GN2026-07-15

LATTICE: Graph Self-Supervised Learning for Multimodal Spatial Omics Integration

Jagan Mohan Reddy Dwarampudi, Veena Kochat, Suresh Satpati, Hien Van Nguyen, Kunal Rai, Tania Banerjee

Spatially resolved omics studies increasingly combine transcriptomic and epigenomic assays, yet downstream analysis is often still performed using single-modality pipelines. We present LATTICE (Latent Alignment of Tissue-level and Transcriptomic Information for Cross-modal Embedd…

View free PDFSource page
arxivq-bio.GNcs.AIcs.LG2026-07-20

Making Single-Cell Data Distillation Auditable: Traceable Real-Cell Coresets via Discrete Min-Max Selection

Yaodi Luo, Peize He, Bowen Han, Lingbei Mengg

Single-cell datasets are increasingly costly to store, audit, and reuse for model training. Dimensionality reduction and dataset distillation can reduce this burden, but conventional distillation methods often produce synthetic expression profiles that cannot be traced to an assa…

View free PDFSource page
arxivcs.LGq-bio.GNq-bio.QM2026-07-06

Data-Driven Soft Labeling Scales DNA Read Classification to Whole-Body Cell-Type Deconvolution

Dmytro Rizdvanetskyi, Nathan Roos, Pavlo Lutsik

Cell-type deconvolution, the task of estimating the proportions of constituent cell types in a heterogeneous biological sample, is a core problem in computational biology. Methods that rely on epigenetic marks such as DNA methylation typically operate on aggregated methylation es…

View free PDFSource page
arxivcs.LGq-bio.GN2026-07-23

HierarchicalDAEW: Domain-Aware Edge-Weighted Graph Convolution with Evidential Uncertainty for Multi-Section Spatial Gene Expression Prediction from H&E Histology

Kritanu Chattopadhyay, Soumya Chatterjee, Ondrej Krejcar, Debotosh Bhattacharjee

Spatial transcriptomics assays remain costly and technically demanding, restricting transcriptome-wide profiling to specialist settings and preventing routine clinical deployment. Predicting spatially resolved gene expression from H&E histology could close this gap, yet current m…

View free PDFSource page
arxivq-bio.GNcs.LG2026-07-15

Screening of Biosecurity Features in Metagenomic Data with Evo 2 Probes

Jeremy Guntoro, Alexander Dack, Dylan Danno, Michaela Jančovičová, Križan Jurinović, Vanessa Smilansky

Genomic foundation models such as Evo 2 learn rich sequence representations, but their value for biosecurity screening is largely unexplored. We ask how much biosecurity-relevant signal is linearly accessible in these representations by training minimal linear and attention probe…

View free PDFSource page