XCC-Net: An X-Shaped Collective Convolution Network Architecture for Medical Image Segmentation
Anass Garbaz, Yassine Oukdach, Said Charfi, Mohamed El Ansari, Lahcen Koutti, Mustapha Hedabou, Mustapha Oujaoura, Abdel Motalib Lagsoun
Encoder–decoder models are widely used for pixel-level segmentation due to their ability to capture and combine multiscale features. However, skip connections between the encoder and decoder often require cropping to mitigate border pixel loss during convolutions, which can introduce inefficiencies and limit performance. This study explores the potential of modifying these connections by removing direct encoder-to-decoder links to enhance segmentation accuracy. We propose a novel architecture, termed XCC-Net, which features two context-capturing pathways and two symmetric pathways for enlargement. These pathways are interconnected via channels, enabling automated detection of structures with varied shapes. The XCC-Net’s X-shaped architecture links skip connections exclusively between encoder-to-encoder and decoder-to-decoder, omitting direct encoder-to-decoder feature transfers to potentially improve performance. The XCC-Net model was evaluated on multiple medical imaging datasets, including wireless capsule endoscopy (WCE), colonoscopy, and dermoscopy images. Experimental results showed that XCC-Net outperformed state-of-the-art segmentation models, achieving dice coefficients of 91.70%, 89.26%, 87.15%, and 79.07% on the MICCAI 2017 (Red Lesion), PH2, CVC-ClinicDB, and ISIC 2017 datasets, respectively. XCC-Net’s X-shaped architecture, with its unique skip connections, demonstrates improved segmentation performance across various medical imaging tasks.