CORTEXA
← Browse
crossrefMachine Learning and Knowledge Extraction2024-04-15Cited by 40

Impact of Nature of Medical Data on Machine and Deep Learning for Imbalanced Datasets: Clinical Validity of SMOTE Is Questionable

Seifollah Gholampour

Dataset imbalances pose a significant challenge to predictive modeling in both medical and financial domains, where conventional strategies, including resampling and algorithmic modifications, often fail to adequately address minority class underrepresentation. This study theoretically and practically investigates how the inherent nature of medical data affects the classification of minority classes. It employs ten machine and deep learning classifiers, ranging from ensemble learners to cost-sensitive algorithms, across comparably sized medical and financial datasets. Despite these efforts, none of the classifiers achieved effective classification of the minority class in the medical dataset, with sensitivity below 5.0% and area under the curve (AUC) below 57.0%. In contrast, the similar classifiers applied to the financial dataset demonstrated strong discriminative power, with overall accuracy exceeding 95.0%, sensitivity over 73.0%, and AUC above 96.0%. This disparity underscores the unpredictable variability inherent in the nature of medical data, as exemplified by the dispersed and homogeneous distribution of the minority class among other classes in principal component analysis (PCA) graphs. The application of the synthetic minority oversampling technique (SMOTE) introduced 62 synthetic patients based on merely 20 original cases, casting doubt on its clinical validity and the representation of real-world patient variability. Furthermore, post-SMOTE feature importance analysis, utilizing SHapley Additive exPlanations (SHAP) and tree-based methods, contradicted established cerebral stroke parameters, further questioning the clinical coherence of synthetic dataset augmentation. These findings call into question the clinical validity of the SMOTE technique and underscore the urgent need for advanced modeling techniques and algorithmic innovations for predicting minority-class outcomes in medical datasets without depending on resampling strategies. This approach underscores the importance of developing methods that are not only theoretically robust but also clinically relevant and applicable to real-world clinical scenarios. Consequently, this study underscores the importance of future research efforts to bridge the gap between theoretical advancements and the practical, clinical applications of models like SMOTE in healthcare.

View free PDFSource page

Related papers

crossrefMachine Learning and Knowledge Extraction2024-12-25Cited by 14

Analyzing the Impact of Data Augmentation on the Explainability of Deep Learning-Based Medical Image Classification

(Freddie) Liu, Gizem Karagoz, Nirvana Meratnia

Deep learning models are widely used for medical image analysis and require large datasets, while sufficient high-quality medical data for training are scarce. Data augmentation has been used to improve the performance of these models. The lack of transparency of complex deep-lea…

View free PDFSource page
crossrefMachine Learning and Knowledge Extraction2025-08-01Cited by 6

Quantum Machine Learning and Deep Learning: Fundamentals, Algorithms, Techniques, and Real-World Applications

Maria Revythi, Georgia Koukiou

Quantum computing, with its foundational principles of superposition and entanglement, has the potential to provide significant quantum advantages, addressing challenges that classical computing may struggle to overcome. As data generation continues to grow exponentially and tech…

View free PDFSource page
crossrefMachine Learning and Knowledge Extraction2023-12-27Cited by 5

Transforming Simulated Data into Experimental Data Using Deep Learning for Vibration-Based Structural Health Monitoring

Abhijeet Kumar, Anirban Guha, Sauvik Banerjee

While machine learning (ML) has been quite successful in the field of structural health monitoring (SHM), its practical implementation has been limited. This is because ML model training requires data containing a variety of distinct instances of damage captured from a real struc…

View free PDFSource page
crossrefMachine Learning and Knowledge Extraction2026-07-22

Alzheimer’s Disease Detection Based on Machine Learning and Deep Learning Frameworks: A Cross-Dataset Comparative Performance Analysis and Assessment of Clinical Readiness

Keenan Ramnarain, Rito Clifford Maswanganyi, Philani Khumalo

Alzheimer’s disease (AD) is the most prevalent neurodegenerative disorder worldwide, affecting approximately 56.9 million people in 2021 and projected to reach 152 million by 2050. Its defining pathological features, amyloid-beta plaques and neurofibrillary tangles, accumulate fo…

View free PDFSource page
crossrefMachine Learning and Knowledge Extraction2025-09-21Cited by 41

Customer Churn Prediction: A Systematic Review of Recent Advances, Trends, and Challenges in Machine Learning and Deep Learning

Mehdi Imani, Majid Joudaki, Ali Beikmohammadi, Hamid Arabnia

Background: Customer churn significantly impacts business revenues. Machine Learning (ML) and Deep Learning (DL) methods are increasingly adopted to predict churn, yet a systematic synthesis of recent advancements is lacking. Objectives: This systematic review evaluates ML and DL…

View free PDFSource page
crossrefMachine Learning and Knowledge Extraction2025-09-02Cited by 2

A Novel Prediction Model for Multimodal Medical Data Based on Graph Neural Networks

Lifeng Zhang, Teng Li, Hongyan Cui, Quan Zhang, Zijie Jiang, Jiadong Li, et al.

Multimodal medical data provides a wide and real basis for disease diagnosis. Computer-aided diagnosis (CAD) powered by artificial intelligence (AI) is becoming increasingly prominent in disease diagnosis. CAD for multimodal medical data requires addressing the issues of data fus…

View free PDFSource page