CORTEXA
← Browse
openalexBMC Medical Informatics and Decision Making2026-07-24Cited by 0

Quality assessment of synthetic data in healthcare: a critical appraisal with a focus on subgroup bias

Christos Chatzichristos, Dimitris Katsimpokis, Congting Lai, L. van Santvliet, Flavio Camarrone, Daan Knoors, Gijs Geleijnse, Michel Van Speybroek, Martine Lewi, Bart Vannieuwenhuyse, Maarten De Vos

The creation of synthetic data in healthcare research has emerged as a compelling solution to address challenges related to data privacy, in case of data-sharing and data scarceness for training machine learning models. The current paper aims to provide a critical landscape analysis of synthetic data quality metrics used in healthcare, mapping the state of the field, examining how current fidelity and utility scores often fail to capture weaknesses that appear in specific patient subgroups, and illustrating these blind spots with a lung-cancer case study. We specifically set out to highlight the need for assessing synthetic datasets not only at the aggregate level but also across clinically meaningful subgroups, to ensure reliability and fairness in downstream applications. By examining the landscape of synthetic data adoption in healthcare, we highlight the methodological and utility considerations that shape its integration into the research ecosystem. We elucidate how synthetic data can expedite research initiatives, support data-driven decision-making, and facilitate innovative methodologies while safeguarding sensitive patient information. Conversely, while many papers focus solely on the advantages of synthetic data, we aim to highlight also the potential constraints of using synthetic data in a real-world case involving lung cancer patient level data. Our analysis centers on identifying situations where synthetic data might introduce biases or inaccuracies that hinder the generation of meaningful clinical insights. We demonstrate the limitations of using synthetic data in a lung cancer case study, especially when dealing with small subgroups from the original dataset. While high values for fidelity and utility metrics are achieved when considering the entire synthetic dataset, significant discrepancies arise in cross-classification performance when examining subsets. Additionally, visual inspection reveals spurious correlations in the synthetic data that are not present in the real data. Global realism scores can give a false sense of security: in our case study, metrics that rate the full dataset as “high quality” overlook errors which appear once we zoom in on specific subgroups. We argue that future work must (i) design subgroup-aware fidelity and utility metrics, (ii) favor conditional generators that model rare strata explicitly, and (iii) report metric panels alongside qualitative, clinical sanity checks. Until such standards mature, our analysis highlights that researchers and regulators should treat synthetic data with caution.

View free PDFSource page

Related papers

openalexBMC Medical Informatics and Decision Making2026-07-24

AI-powered spectral CT analysis for clinical decision support in carotid vulnerable plaque detection: a deep learning approach

Yunzhe Ni, Tianyu Zhang, Zonghui Huang, Yue Wang, Guochao Han, Lin Yuan, et al.

Carotid vulnerable plaques (CVPs) represent a major cause of ischemic stroke, yet current diagnostic methods lack sufficient precision for early detection. Spectral computed tomography (CT) enables detailed plaque characterization, but its clinical utility depends on advanced ana…

View free PDFSource page
openalexBMC Medical Informatics and Decision Making2026-07-24

AI-powered risk prediction models for preventable maternal mortality in rural settings: a systematic review

Joy Aifuobhokhan, Ayodeji Ogunjinmi, Chukwuemeka Abraham Agbarakwe, Deborah Oladunmolu Oduguwa, Annie Peter Essiet, Temitayo Osunkiyesi, et al.

Maternal mortality remains disproportionately high in low- and middle-income countries, particularly in rural settings with limited access to skilled obstetric care. Artificial intelligence and machine learning models offer promise for early risk prediction, yet their methodologi…

View free PDFSource page
openalexBMC Medical Informatics and Decision Making2026-07-24

Development and temporal validation of a machine learning-based model for predicting renal impairment in multiple myeloma: a single-center retrospective cohort study

Manli Zhou, Sisi Feng

Multiple myeloma (MM) is a hematopoietic system malignancy characterized by clonal proliferation of abnormal plasma cells, commonly presenting with renal impairment (RI) that significantly impacts patient’s quality of life. The objective of this study was to develop a predictive…

View free PDFSource page