CORTEXA
← Browse
arxivcs.CV2026-07-09

CT-CLIP Representations for Multimodal Lung Cancer Survival Prediction

Sofie Allgöwer, Mikael Johansson, Andreas Hallqvist, Jonas Andersson, Åse Johnsson, Ida Häggström, Jennifer Alvén

Accurate prognosis prediction is important for treatment planning in lung cancer, but deep learning-driven survival modelling is often limited by the scarcity of curated imaging cohorts with reliable outcome data. This study evaluates whether representations from a domain-specific foundation model can be used for multimodal survival prediction in data-constrained clinical settings. We assess the foundation model CT-CLIP as a feature extractor for pretreatment computed tomography images and clinical variables from 242 diagnosed lung cancer patients. The evaluation includes adaptation strategies based on frozen encoders, full fine-tuning, and low-rank adaptation, together with modality ablations and comparisons with clinical and multimodal baselines. The results show that a frozen CT-CLIP model combined with a trainable lightweight survival head outperforms the clinical baseline and achieves comparable or improved performance relative to other multimodal approaches, and separates patients into clinically meaningful high- and low-risk groups.

View free PDFSource page

Related papers

arxivcs.LGcs.AIcs.CV2026-06-28

AdaSurvMamba: Dynamic Fusion and Semantic Scanning for Multimodal Survival Analysis

Jialong Zhong, Tingwei Liu, Baokun Yue, Jingjing Li, Yongri Piao, Miao Zhang, et al.

Multimodal survival analysis utilizing whole slide images (WSIs) and genomic profiles is fundamental for cancer prognosis. Recently, state-space models like Mamba have emerged as powerful tools for sequence modeling. However, translating this success to complex multimodal tasks i…

View free PDFSource page
arxivcs.CVcs.LG2026-07-16

AlphaWiSE: Adaptive Weight Interpolation for Continual Multimodal Representation Learning

Sarthak Jain, Qiran Hu, Zhen Zhu, Yaoyao Liu

Multimodal models such as CLIP learn a shared embedding space for cross-modal retrieval, but continual adaptation to sequentially arriving data can disrupt the cross-modal alignment acquired from earlier phases. Conventional continual-learning methods return a single checkpoint,…

View free PDFSource page
arxivcs.CV2026-07-24

Twins: Learn to Predict Unified Representations with Focal Loss

Kaixiong Gong, Xin Cai, Bin Lin, Hao Wang, Yunlong Lin, Mingzhe Zheng, et al.

Unified multimodal models seek a shared visual token space that supports both multimodal understanding and image generation. Discrete methods unify the interface via a shared codebook, whereas continuous pipelines often rely on two disparate representations -- semantic features (…

View free PDFSource page
arxivcs.CV2026-06-28

When Does Synthetic CT Transfer? A Label-Free Donor/Host Diagnostic for Medical Vision-Language Model Routing on Real Lung CT

Fakrul Islam Tushar

A synthetic measurement of model competence is useful only if it survives the move to real data, yet the real labels that would verify it are exactly what medical imaging lacks. We ask whether transfer can be predicted in advance, label-free, and answer with a mechanism: on synth…

View free PDFSource page
arxivcs.CV2026-07-15

TRACE-PCa: Predicting Prostate Cancer Progression from Longitudinal MRI During Active Surveillance

Hongye Zeng, Shreeram Athreya, Dingyuan Dai, Steve Raman, Leonard Marks, William Speier, et al.

Active surveillance (AS) is the preferred strategy for favorable-risk prostate cancer, yet current protocols rely on scheduled repeat biopsies, most of which reveal no progression and are unnecessary. Existing risk-stratification tools operate on single time-point imaging or depe…

View free PDFSource page