CORTEXA
← Browse
openalexApplied Artificial Intelligence2026-07-24Cited by 0

Deep Spectrogram Learning for Intelligent Classification of Respiratory Disorders

Vedant Agnihotri, S. Vidivelli, Momina Shaheen, R. Manikandan, S. Magesh

Such conditions as chronic obstructive pulmonary disease (COPD), pneumonia, and upper respiratory tract infections (URTI) form a significant health concern worldwide, which is why the development of effective diagnostic methods to detect them early deserves attention. The current study presents a deep learning model devoted to classifying audios of respiratory sounds – sourced from the ICBHI Respiratory Sound Database – into one of four classes: normal, crackles, wheezes, and combinations of crackles and wheezes. Preprocessing of the audio data involved resampling and segmentation, followed by the use of Short-Time Fourier Transform (STFT) to produce spectrograms. A set of data augmentation techniques, including time stretching, pitch shifting, and adding noise, was applied to expand the dataset. Unlike prior work, our study introduces a lightweight yet high-performing spectrogram-based CNN framework (ResNet18 and ResNet50) that is validated under patient-level cross-validation using the ICBHI dataset, addressing real-world deployability and generalization challenges in clinical settings. The imbalance between classes was handled using Focal Loss, which significantly improved the detection of minority class patterns such as combined crackles and wheezes. The ResNet18 and ResNet50 were trained with GroupKFold cross-validation to ensure strict patient separation. The ResNet50 achieved the best average harmonic mean score of 0.8429 and specificity of 0.9232.

View free PDFSource page