CORTEXA
← Browse
crossrefJournal of Imaging2025-07-18Cited by 2

A Novel 3D Convolutional Neural Network-Based Deep Learning Model for Spatiotemporal Feature Mapping for Video Analysis: Feasibility Study for Gastrointestinal Endoscopic Video Classification

Mrinal Kanti Dhar, Mou Deb, Poonguzhali Elangovan, Keerthy Gopalakrishnan, Divyanshi Sood, Avneet Kaur, Charmy Parikh, Swetha Rapolu, Gianeshwaree Alias Rachna Panjwani, Rabiah Aslam Ansari, Naghmeh Asadimanesh, Shiva Sankari Karuppiah, Scott A. Helgeson, Venkata S. Akshintala, Shivaram P. Arunachalam

Accurate analysis of medical videos remains a major challenge in deep learning (DL) due to the need for effective spatiotemporal feature mapping that captures both spatial detail and temporal dynamics. Despite advances in DL, most existing models in medical AI focus on static images, overlooking critical temporal cues present in video data. To bridge this gap, a novel DL-based framework is proposed for spatiotemporal feature extraction from medical video sequences. As a feasibility use case, this study focuses on gastrointestinal (GI) endoscopic video classification. A 3D convolutional neural network (CNN) is developed to classify upper and lower GI endoscopic videos using the hyperKvasir dataset, which contains 314 lower and 60 upper GI videos. To address data imbalance, 60 matched pairs of videos are randomly selected across 20 experimental runs. Videos are resized to 224 × 224, and the 3D CNN captures spatiotemporal information. A 3D version of the parallel spatial and channel squeeze-and-excitation (P-scSE) is implemented, and a new block called the residual with parallel attention (RPA) block is proposed by combining P-scSE3D with a residual block. To reduce computational complexity, a (2 + 1)D convolution is used in place of full 3D convolution. The model achieves an average accuracy of 0.933, precision of 0.932, recall of 0.944, F1-score of 0.935, and AUC of 0.933. It is also observed that the integration of P-scSE3D increased the F1-score by 7%. This preliminary work opens avenues for exploring various GI endoscopic video-based prospective studies.

View free PDFSource page

Related papers

crossrefJournal of Imaging2026-07-24

Organ Segmentation with Machine Learning Models

Alexandros Barmperis, Olga Menegaki, Anna Panagiotakopoulou, Andreas Vezakis, Ioannis Vezakis, Ioannis Kakkos, et al.

Accurate segmentation of abdominal organs in Computed Tomography (CT) underpins radiotherapy planning, surgical planning, and disease monitoring. Existing benchmarks rank architectures by a single aggregate Dice score, without per-organ statistical testing or boundary-sensitive m…

View free PDFSource page
crossrefJournal of Imaging2026-07-08

Deep Learning-Based Multi-Class Pediatric Wrist Fracture Subtype Classification: A Pilot Study Comparing Convolutional Neural Network Architectures

Rohan A. Phadke, Samer G. Salman, Zane G. Salman, Sai M. Yedupati, Joshua Ong, Alireza Tavakkoli, et al.

Pediatric wrist fractures are among the most prevalent musculoskeletal injuries in children. Fracture subtype, including buckle/torus, greenstick, and Salter–Harris physeal injuries, directly influences management and prognosis. Subspecialty radiographic expertise required for su…

View free PDFSource page
crossrefJournal of Imaging2026-03-16

Advanced Sensitive Feature Machine Learning for Aesthetic Evaluation Prediction of Industrial Products

Jinyan Ouyang, Ziyuan Xi, Jianning Su, Shutao Zhang, Ying Hu, Aimin Zhou

As product aesthetics increasingly drive consumer preference, quantitative evaluation remains hindered by subjective evaluation biases and the black-box nature of modern artificial intelligence. This study proposes an advanced machine learning framework incorporating sensitivity-…

View free PDFSource page
crossrefJournal of Imaging2026-02-18Cited by 5

Analysis of Biological Images and Quantitative Monitoring Using Deep Learning and Computer Vision

Aaron Gálvez-Salido, Francisca Robles, Rodrigo J. Gonçalves, Roberto de la Herrán, Carmelo Ruiz Rejón, Rafael Navajas-Pérez

Automated biological counting is essential for scaling wildlife monitoring and biodiversity assessments, as manual processing currently limits analytical effort and scalability. This review evaluates the integration of deep learning and computer vision across diverse acquisition…

View free PDFSource page
crossrefJournal of Imaging2025-10-09Cited by 1

Non-Destructive Volume Estimation of Oranges for Factory Quality Control Using Computer Vision and Ensemble Machine Learning

Wattanapong Kurdthongmee, Arsanchai Sukkuea

A crucial task in industrial quality control, especially in the food and agriculture sectors, is the quick and precise estimation of an object’s volume. This study combines cutting-edge machine learning and computer vision techniques to provide a comprehensive, non-destructive me…

View free PDFSource page
crossrefJournal of Imaging2025-09-23Cited by 38

A Review on the Detection of Plant Disease Using Machine Learning and Deep Learning Approaches

Thandiwe Nyawose, Rito Clifford Maswanganyi, Philani Khumalo

The early and accurate detection of plant diseases is essential for ensuring food security, enhancing crop yields, and facilitating precision agriculture. Manual methods are labour-intensive and prone to error, especially under varying environmental conditions. Artificial intelli…

View free PDFSource page