CORTEXA
← Browse
openalexZenodo (CERN European Organization for Nuclear Research)2026-07-24Cited by 0

Federated Machine Learning for Privacy-Preserving Healthcare Risk Prediction

Jnana Durga Ravi Chand Mutthina, Sriharsha Doniparthi, Ramya Kata

This preprint presents an Explainable Federated Machine Learning framework, called XFML, for privacy-preserving diabetes risk prediction across distributed healthcare institutions. The framework enables multiple hospital nodes to collaboratively develop a predictive model without transferring or centralizing sensitive patient-level data. XFML combines locally trained XGBoost and Random Forest models with a federated stacking meta-learner, adaptive variance-penalized feature importance aggregation, Gaussian differential privacy, and dual-layer explainability using SHAP and LIME. The framework is evaluated using 445,132 records from the 2022 CDC Behavioral Risk Factor Surveillance System, partitioned across five geographically distributed, non-IID hospital nodes. The proposed model achieves an AUC-ROC of 0.8247 and sensitivity of 0.8249, with a minimal performance difference of −0.0005 AUC compared with the centralized baseline. It also provides improved probability calibration, achieving a Brier score of 0.1077 compared with 0.1748 for the centralized model. Explainability analysis identifies general health, age, and body mass index as the most influential diabetes-risk factors, while stability experiments demonstrate highly consistent federated SHAP explanations. The findings demonstrate that privacy preservation, predictive performance, probability calibration, and model interpretability can be jointly achieved in federated healthcare machine learning. This manuscript is deposited as a preprint and has not yet completed formal peer review.

View free PDFSource page

Related papers

openalexZenodo (CERN European Organization for Nuclear Research)2026-07-26

AI-Driven Intrusion Detection for the Internet of Things: A Scoping Review of Federated Learning, Privacy-Preserving Architectures, and Edge Deployability

Gilbert Aimufua, Godwin Agbonkhese

Federated learning has emerged as the dominant architectural response to the privacy and communication constraints of centralised intrusion detection in Internet of Things environments, yet the field lacks a synthesis that maps the concurrent state of architecture diversity, priv…

View free PDFSource page
openalexZenodo (CERN European Organization for Nuclear Research)2026-07-26

Interpretable machine-learning risk stratification at diagnosis for 3-year mortality in de novo metastatic prostate cancer (SEER): reproducibility code

Xin Wang, Guanglei Yao, Wei Ding

This archive contains the analysis code, the predictor dictionary, and the retrained primary model objects underlying the manuscript "Interpretable machine-learning risk stratification at the time of diagnosis for 3-year mortality in de novo metastatic prostate cancer: developmen…

View free PDFSource page
openalexZenodo (CERN European Organization for Nuclear Research)2026-08-09

A Systematic Review of Machine Learning, Deep Learning, and Explainable AI Approaches for Cardiac Disease Prediction

Sunanda Budihal, Sheetalrani Kawale, Abhishek Angadi

The cardiovascular (Cardiac) disease (CVD) is another factor that causes death among the global population most, and this is the reason why there is a high necessity to implement proper, effective, and interpretive diagnostic systems. The usage of machine learning (ML), deep lear…

Also available via: European Organization for Nuclear Research

View free PDFSource page
openalexZenodo (CERN European Organization for Nuclear Research)

Automatic classification of cyber incidents using privacy-preserving artificial intelligence

Loya Caroldene Haughton, Eduardo Fidalgo, David Lewis

As cyber incidents increase in complexity, diversity and frequency, cybersecurity practitioners find it more challenging to extract meaningful threat intelligence and insights from cyber incident reports. This problem is worsened by the limited number of these reports; as cyber i…

Also available via: European Organization for Nuclear Research

View free PDFSource page
openalexZenodo (CERN European Organization for Nuclear Research)2026-07-24

Before the Model: Why Datasets and Data Representation Define What Machine Learning Can Learn

Jean Franck Loa Rojas

Machine learning systems do not learn reality directly; they learn from the representations preserved in their datasets. This structured narrative review examines how dataset purpose, coverage, integrity, labeling, independence, reproducibility, governance, and continuity determi…

View free PDFSource page