Unmasking adversarial attacks using a robust XAI-driven approach for secure medical image classification
Sudarshan Saha, Md. Sohel Rana, Ahmed Wasif Reza, M. R. Amin
Deep learning models have revolutionized medical image analysis but remain vulnerable to adversarial attacks. This work presents FusionStack, an explainability-driven framework for detecting adversarial perturbations through multi-method XAI stability analysis. We introduce three novel stability metrics (LIME Stability Metric, SHAP Stability Metric, Grad-CAM Stability Metric) quantifying explanation consistency across convolutional feature hierarchies, augmented with frequency-domain and gradient-based features. Tested on five medical imaging datasets, five CNN models, and five adversarial attacks with rigorous patient-level separation, FusionStack achieves 98.71% accuracy and 0.998 AUC. Cross-attack evaluation shows DeepFool-trained detectors achieve 96.85% transferability as universal defenders. Cross-dataset evaluation demonstrates 87.05% mean accuracy, with histopathology-trained detectors achieving 100% transfer to all modalities. Lightweight frequency-domain detection (10 features) achieves 95.46% accuracy with $$5.75\times$$ speedup. Our findings establish XAI stability analysis as a principled pathway toward adversarially robust medical AI systems for safety-critical applications.