Early diagnosis and risk stratification of aortic stenosis using artificial intelligence applied to echocardiography: scoping review
Gabriel Eduardo Malagón Tarqui, Juanita Valencia García, Erwin Hernando Hernández Rincón
TL;DR: Artificial intelligence algorithms perform well in detecting and classifying the severity of aortic stenosis, and shows high diagnostic potential in retrospective datasets, reaching metrics that emulate expert accuracy.
INTRODUCTION Aortic stenosis (AS) is the most common acquired valvular heart disease worldwide, accounting for 43% of valvular diseases. It is estimated that 40-50% of patients with severe symptomatic AS do not receive intervention, resulting in a mortality rate of over 90%. Transthoracic echocardiography remains the gold standard for diagnosis, making it critical for early detection of the disease. However, it is operator-dependent and varies according to the patient's clinical presentation. In this context, artificial intelligence algorithms, especially deep learning algorithms applied to echocardiography, are emerging as tools with the potential to automate and improve the detection of aortic stenosis. OBJECTIVE To evaluate the available evidence on the usefulness of artificial intelligence tools applied to echocardiography for the early diagnosis of aortic stenosis, identifying their performance, clinical applicability, and methodological limitations. METHODS A scoping review was conducted in four databases (PubMed, Scopus, Web of Science, and BIREME) in accordance with the PRISMA-ScR guideline, which included 25 studies between January 2020 and December 2025 that used AI systems applied to echocardiography for the early diagnosis and risk stratification of aortic stenosis. RESULTS Twenty-five studies met the inclusion criteria for this review. Artificial intelligence (AI) algorithms, especially convolutional neural networks, achieved heterogeneous performance. The AUC ranged from 0.82 to 0.99; sensitivity was 82.2-90% and specificity was 88-99%. Multivision models performed better than single-vision models. CONCLUSIONS Artificial intelligence algorithms perform well in detecting and classifying the severity of AS. Their performance shows high diagnostic potential in retrospective datasets, reaching metrics that emulate expert accuracy. Critical barriers remain, such as lack of external validation, interpretability, and clinical integration. Prospective multicenter studies with harmonized regulatory frameworks are needed for global validation.