CORTEXA
← Browse
arxivq-bio.GNcs.AI2026-06-25

GRAFT: Biological Graph and Hypergraph Benchmarks for Linked Gene Expression and Phenotypic Trait Prediction in Arabidopsis thaliana

Manuel Serna-Aguilera, Vanshika Jindal, Fiona L. Goggin, Jiamei Li, Aranyak Goswami, Alexander Bucksch, Suxing Liu, Khoa Luu

Understanding which genes control which traits in an organism remains one of the central challenges in biology. Despite significant advances in data collection technology, our ability to map genes to traits is still limited. This genome-to-phenome (G2P) challenge spans several problem domains, including plant breeding, and requires methods capable of reasoning over high-dimensional, heterogeneous, and biologically structured data. Current datasets and data repositories, however, are not well-equipped for this task. Current studies do not link gene expression and trait data, and most focus on very specific traits, limiting the breadth of possible correlations. To address this gap, we present the novel Gene-Graph Regression for Arabidopsis Functional Traits (GRAFT) dataset, a curated multi-modal dataset linking gene expression profiles with phenotypic trait measurements in Arabidopsis thaliana, a model organism in plant biology. GRAFT supports tasks such as phenotype prediction and interpretable graph learning. In addition, we benchmark conventional regression and explanatory baselines, including a biologically-informed hypergraph baseline, to validate gene-trait associations. To the best of our knowledge, this is the first dataset to provide multimodal gene information and heterogeneous trait or phenotype data for the same Arabidopsis thaliana specimens. With GRAFT, we aim to foster research to accurately understand the relationship between genotypes and phenotypes using gene information, higher-order gene pairings, and trait data from multiple sources.

View free PDFSource page

Related papers

arxivq-bio.GNcs.AI2026-07-22

Foundation-model-guided radiogenomic discovery linking cancer genomes to cancer scans

Frederik Hauke, Jeremias Krause, Patrick Wienholt, Christiane Kuhl, Ingo Kurth, Sikander Hayat, et al.

The function of many genes is still unknown, and conventional driver-discovery methods, which rely on how frequently a gene is mutated, cannot assess genes that are only rarely affected. Here we pair Evo~2-based genome analysis with routine clinical imaging to identify gene--phen…

View free PDFSource page
arxiveess.IVcs.AIq-bio.GN2026-06-29

Data-Efficient Multimodal Alignment for Histopathology-based Molecular Prediction

Dominik Winter, Dominik Vonficht, Loïc Le Bescond, Christian Gebbe, Marco Rosati, Richard J. Chen, et al.

H&E-stained whole-slide images offer cohort-scale availability and rich spatial context but lack molecular specificity, whereas bulk RNA-seq provides transcriptome-wide resolution at high cost with limited archival availability. We show that training a lightweight alignment modul…

View free PDFSource page
arxivq-bio.GNcs.AIcs.LG2026-07-20

Making Single-Cell Data Distillation Auditable: Traceable Real-Cell Coresets via Discrete Min-Max Selection

Yaodi Luo, Peize He, Bowen Han, Lingbei Mengg

Single-cell datasets are increasingly costly to store, audit, and reuse for model training. Dimensionality reduction and dataset distillation can reduce this burden, but conventional distillation methods often produce synthetic expression profiles that cannot be traced to an assa…

View free PDFSource page
arxivcs.LGcs.AIq-bio.GN2026-07-04

SHIFT: Survival Prediction from Incomplete and Heterogeneous Genomic Data

Muhammet Sami Yavuz, Ayhan Can Erdur, Sabri Mustafa Kahya, Benedikt Wiestler, Jana Lipkova

Genomic prediction models often fail to transfer across institutions because sequencing panels differ across sites, creating structural feature missingness at deployment. Existing approaches to this challenge typically restrict analysis to genes shared across cohorts, exclude pat…

View free PDFSource page
arxivcs.LGcs.AIq-bio.BMq-bio.GN2026-06-26

Two-Stage Fine-Tuning for Protein Sequence Generation with Targeted Amino-Acid Composition

Violeta Basten-Romero, Rubén Muñoz-Tafalla, Anna María Díaz-Rovira, Bertran Miquel-Oliver, Isaac Filella-Merce, Víctor Guallar

Protein language models are standard priors for biological sequence generation, but steering them toward explicit distributional design targets remains largely unexplored. We study a constrained protein generation problem in which sequences must match a desired amino-acid (AA) co…

View free PDFSource page
arxivq-bio.GNcs.AI2026-06-26

Reconstructing the Developmental Trajectory of Adipocytes in Human Adipose Tissue Using Single-Cell RNA Sequencing

Weny S. M Sitinjak, Humasak Tommy Argo Simanjuntak

Obesity is a global health crisis associated with metabolic disorders such as type 2 diabetes and cardiovascular disease. This study employed single-cell RNA sequencing to reconstruct the developmental trajectory of human adipocytes from adipose tissue samples. Our analysis ident…

View free PDFSource page