Quality assessment of synthetic data in healthcare: a critical appraisal with a focus on subgroup bias
Christos Chatzichristos, Dimitris Katsimpokis, Congting Lai, L. van Santvliet, Flavio Camarrone, Daan Knoors, Gijs Geleijnse, Michel Van Speybroek, Martine Lewi, Bart Vannieuwenhuyse, Maarten De Vos
The creation of synthetic data in healthcare research has emerged as a compelling solution to address challenges related to data privacy, in case of data-sharing and data scarceness for training machine learning models. The current paper aims to provide a critical landscape analysis of synthetic data quality metrics used in healthcare, mapping the state of the field, examining how current fidelity and utility scores often fail to capture weaknesses that appear in specific patient subgroups, and illustrating these blind spots with a lung-cancer case study. We specifically set out to highlight the need for assessing synthetic datasets not only at the aggregate level but also across clinically meaningful subgroups, to ensure reliability and fairness in downstream applications. By examining the landscape of synthetic data adoption in healthcare, we highlight the methodological and utility considerations that shape its integration into the research ecosystem. We elucidate how synthetic data can expedite research initiatives, support data-driven decision-making, and facilitate innovative methodologies while safeguarding sensitive patient information. Conversely, while many papers focus solely on the advantages of synthetic data, we aim to highlight also the potential constraints of using synthetic data in a real-world case involving lung cancer patient level data. Our analysis centers on identifying situations where synthetic data might introduce biases or inaccuracies that hinder the generation of meaningful clinical insights. We demonstrate the limitations of using synthetic data in a lung cancer case study, especially when dealing with small subgroups from the original dataset. While high values for fidelity and utility metrics are achieved when considering the entire synthetic dataset, significant discrepancies arise in cross-classification performance when examining subsets. Additionally, visual inspection reveals spurious correlations in the synthetic data that are not present in the real data. Global realism scores can give a false sense of security: in our case study, metrics that rate the full dataset as “high quality” overlook errors which appear once we zoom in on specific subgroups. We argue that future work must (i) design subgroup-aware fidelity and utility metrics, (ii) favor conditional generators that model rare strata explicitly, and (iii) report metric panels alongside qualitative, clinical sanity checks. Until such standards mature, our analysis highlights that researchers and regulators should treat synthetic data with caution.