Reproducibility Package for Explainable and Leakage-Conscious Machine Learning for Athlete Injury Risk Modeling Across Heterogeneous Datasets
Abdülkadir Enes GÖRGÜLÜ, Eray Dursun, Serdar SOLAK
This reproducibility package supports the manuscript “Explainable and Leakage-Conscious Machine Learning for Athlete Injury Risk Modeling Across Heterogeneous Datasets.” It contains the executed and clean analysis notebooks, the corresponding Python script, exact software-version and run-configuration records, complete model hyperparameters, dataset provenance and leakage-audit files, locked figures and tables, machine-readable predictions, calibration outputs, cluster-bootstrap summaries, sensitivity analyses, injury-episode evaluation outputs, explanation-stability results, trained model files, and file-level checksums. The workflow evaluates two externally supplied three-class risk-label datasets and the longitudinal SoccerMon next-14-day injury-episode-start task using group-disjoint or temporal validation, non-informative baselines, ordinal metrics, calibration, uncertainty estimation, event-level analysis, permutation importance, and SHAP-based model interrogation. Dataset files are included only where their applicable licences permit redistribution; otherwise, source identifiers, retrieval information, retained-file characteristics, and SHA-256 checksums are supplied for provenance and verification.