VEA-net: vascular enhancement attention with dual-backbone multi-task learning for comprehensive ROP management across multi-center datasets
Yuehua Chen, Jiajun Wan, Tian Zhang, Zhijiang Wan, Jin Hong, Youpeng Jin
Background Retinopathy of Prematurity (ROP) is a vasoproliferative retinal disorder and a major cause of childhood blindness worldwide. Clinical ROP management involves three correlated tasks: Plus disease detection, stage classification, and treatment decision-making. Although deep learning (DL) has shown promise for automated ROP management, several challenges remain: (1) retinal vascular morphology is central to Plus disease assessment, but DL-based vascular feature modeling is not always explicitly integrated into unified end-to-end multi-task frameworks; (2) Plus disease detection, stage classification, and treatment decision-making are clinically related, yet they are often modeled as separate tasks; and (3) model robustness may be affected by class imbalance and domain shifts across heterogeneous datasets. Methods We propose VEA-Net, a novel dual-backbone multi-task learning framework that addresses these challenges through three core components: (1) We introduce the Vascular Enhancement Attention (VEA) module, which explicitly models and enhances vascular features through multi-scale convolutions, directional selective filtering, and dual attention mechanisms. (2) We develop a hierarchical multi-task learning architecture that jointly optimizes Plus detection, stage classification, and treatment decision-making while leveraging task correlations through hierarchical consistency losses. (3) We implement a dual-domain adaptation strategy combining Domain-Adversarial Neural Networks (DANN) with Maximum Mean Discrepancy (MMD) to learn domain-invariant representations across heterogeneous data sources. Results We validate VEA-Net on three public ROP datasets (FARFUM-ROP, Ostrava, and A-Fundus), comprising 8,636 retinal images from different geographic regions and imaging settings. Using subject-level 10-fold cross-validation with Group K-Fold splitting, our model achieves test AUROC of 91.76% <mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="m1"> <mml:mrow> <mml:mo>±</mml:mo> </mml:mrow> </mml:math> 4.16% for Plus detection, 88.26% <mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="m2"> <mml:mrow> <mml:mo>±</mml:mo> </mml:mrow> </mml:math> 4.61% for stage classification, and 66.81% <mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="m3"> <mml:mrow> <mml:mo>±</mml:mo> </mml:mrow> </mml:math> 15.87% for treatment decision-making. Conclusion VEA-Net provides a unified framework for image-based ROP management by integrating vascular enhancement, hierarchical multi-task learning, and domain adaptation. The model supports Plus detection, stage classification, and auxiliary treatment decision support within a single architecture. The results indicate VEA-Net can learn robust representations across heterogeneous public datasets, and potential for improving ROP management efficiency.