Anonymized but Useful Synthetic Tabular Health Data for AI based Fall Risk Assessment
Ivana Nanevski, Sebastian Jäger, Maryam Mohebi, Matthias Schulte-Althoff, Jörg Pohle, Nicholas Chandler, Rahel Gubser, Alessia Nowak, Fabian Praßer, Daniel Fürstenau, Felix Balzer, Felix Biessmann
Abstract Artificial Intelligence (AI) bears potential for improving health care, but this depends on the availability of open-access, realistic, and useful data. To facilitate AI model development in health care we release SynTabFall, a novel synthetic dataset for fall risk assessment. With a total of 745,380 samples and 44 attributes such as demographics, diseases, mobility and cognition related risk factors, this tabular dataset allows for training fall risk prediction models without access to the original patient data. Models trained on our synthetic dataset can reach predictive performance scores in fall risk assessment which are on par with models trained on real data. To support others in sharing data we also describe a process that was developed over multiple years in one of Germany’s largest hospitals in close collaboration between data protection officers, health care staff, informaticians and AI engineers. The proposed data sharing approach combines established methods for anonymization and modern generative AI (genAI) methods for synthesizing tabular data and allows for sharing health care data responsibly without sacrificing its utility. We release the synthetic fall risk dataset along with the software developed for synthetic data generation and evaluation.