CORTEXA
← Browse
arxivcs.LGq-bio.QM2026-07-10

Variable-Length Generative Protein Design via Generalized Poisson Flow

Chaoran Cheng, Zhanghan Ni, Yanru Qu, Yuxin Chen, Ruihan Guo, Jiajun Fan, Ge Liu

The ability to generate variable-length proteins is crucial in protein design, where the optimal length is often unknown and tightly coupled to designability. Current diffusion- and flow-based generative models typically require the protein length to be specified before sampling, limiting their flexibility in exploring the feasible design space. To address this limitation, we introduce Generalized Poisson Flow (GPFlow), a variable-length generative framework that learns the rate function of an inhomogeneous generalized Poisson process by minimizing its negative log-likelihood. We establish population-level guarantees for recovering the joint multimodal distribution and derive an upper bound on the KL divergence between the data and generated distributions. We comprehensively evaluate GPFlow across structure and sequence design, motif scaffolding, and peptide co-design, spanning Euclidean, categorical, and Riemannian modalities to fully validate its variable-length generation quality. In unconditional design, GPFlow improves structural designability and achieves the best distributional fitness for sequence design compared to their corresponding fixed-length baselines, while perfectly recovering the length distribution. In conditional motif scaffolding, GPFlow ranks first on 10 of 16 structure-based design tasks with significantly more unique successes and also achieves more passed tasks in sequence-based design. In peptide co-design, GPFlow remains competitive even without access to a native-length oracle.

View free PDFSource page

Related papers

arxivcs.LGcs.DCq-bio.QM2026-07-03

Design-CP: Context Parallelism for Design of Protein Nanoparticles

Lorenzo Tarricone, Helen E. Eisenach, Aiko Muraishi, Charlotte M. Deane

Many all-atom generative protein models can in principle design large multimeric complexes by jointly modelling all chains, but their quadratic token- and atom-pair representations quickly exceed single-GPU memory as the number of chains and residues modelled grows. We introduce…

View free PDFSource page
arxivq-bio.QMcs.LG2026-07-15

DyneTrion: A Spatio-temporally Coherent Generative Emulator for Protein Dynamics Across Timescales

Kaihui Cheng, Zhiqiang Cai, Peng Tu, Yisong Yao, Limei Han, Libo Wu, et al.

Proteins function through coordinated motion across multiple spatial and temporal scales, underpinning processes such as ligand binding, allostery, and catalysis. However, accessing long-timescale conformational change through molecular dynamics (MD) simulations remains prohibiti…

View free PDFSource page
arxivq-bio.QMcs.AIcs.LG2026-07-09

DrugGen 2: A disease-aware language model for enhancing drug discovery

Ali Motahharynia, Mohammadreza Ghaffarzadeh-Esfahani, Mahsa Sheikholeslami, Navid Mazrouei, Matin Irajpour, Yousof Gheisari, et al.

Current computational approaches for drug design typically focus on generating molecules conditioned on specific targets or general molecular properties, often neglecting the influence of disease context on target behavior and therapeutic outcomes. To address this gap, we introdu…

View free PDFSource page
arxiveess.AScs.AIcs.LGq-bio.QM2026-06-26

Listening Between the Lines: Joint Learning of ASR Embeddings and LLM-Augmented Linguistics for Dementia Detection

Olivier Jiyoun Jung, Jonghyeon Park, Myungwoo Oh

Early detection of dementia through speech analysis offers a non-invasive screening alternative, but capturing both acoustic and linguistic biomarkers remains challenging. We propose a multimodal framework leveraging Whisper for dual-purpose extraction: acoustic representations f…

View free PDFSource page
arxivcs.LGcs.CEq-bio.QM2026-06-29

Preference-based Antibody Expression Ranking: Scaling with Large-scale Weak Supervision

Josh Qixuan Sun, Morteza Babaie, Wenyang Hou, Mark Crowley, David Young

Antibody expression ranking is a critical task in antibody design, yet its modelling is severely hindered by the scarcity of labeled experimental data. To address this, we propose a unified preference-based learning framework that integrates scarce quantitative expression data wi…

View free PDFSource page
arxivcs.LGq-bio.QMstat.ML2026-06-30

Can Tabular In-Context Learners Generalize to Biomolecular Property Prediction?

Davy Guan, Lu Zhang, Asiri Wijesinghe, Allen Zhu, He Zhao, Helen Power, et al.

Predicting biomolecular properties from limited labeled data is a central bottleneck in protein engineering and small-molecule design. As strong pretrained encoders now supply rich fixed-length representations, the difficulty has shifted from representation learning to building a…

View free PDFSource page