CORTEXA
← Browse
openalexZenodo (CERN European Organization for Nuclear Research)2026-07-26Cited by 0

Reproducibility package for Explainable and Leakage-Conscious Machine Learning for Supplied Injury-Risk Classification and Longitudinal Athlete Injury Forecasting

Abdülkadir Enes GÖRGÜLÜ, Eray Dursun, Serdar Solak

This record provides the complete reproducibility package for the manuscript “Explainable and Leakage-Conscious Machine Learning for Supplied Injury-Risk Classification and Longitudinal Athlete Injury Forecasting.” Overview The study evaluates explainable and leakage-conscious machine-learning workflows across three distinct analytical settings. Two supplied-label datasets are used to examine low-, medium-, and high-risk classification under heterogeneous feature structures, while the longitudinal SoccerMon analysis evaluates future injury-event forecasting from temporally ordered athlete-monitoring data. Analytical scope The package supports: • supplied-label injury-risk classification using a personalized sports-health and training dataset;• supplied-label injury-risk classification using a biomechanical injury-prevention dataset; and• longitudinal prediction of new injury episodes within a 14-day follow-up window using the subjective monitoring component of SoccerMon. Because the provenance, participant characteristics, real-versus-synthetic status, and original target-generation procedures of the two supplied-label datasets could not be independently verified, these datasets are treated as methodological stress tests rather than as prospective or independently validated injury cohorts. SoccerMon serves as the primary longitudinal future-event forecasting case. Validation and leakage-control framework The analytical workflow includes: • group-disjoint hold-out testing and repeated stratified group cross-validation for repeated-user data;• repeated stratified cross-validation for the biomechanical supplied-label task;• temporally ordered 2020-to-2021 validation for SoccerMon;• evaluation within a continuing-athlete panel;• training-partition-only preprocessing, imputation, encoding, scaling, model fitting, calibration, and threshold selection;• exclusion of identifiers, direct target encodings, recommendation fields, inappropriate timestamps, and outcome-derived variables;• dummy-classifier and candidate-model comparisons;• calibration assessment and training-only recalibration;• cluster-bootstrap uncertainty estimation at the user, subject, or athlete level;• sensitivity analyses and feature-group ablation;• permutation importance and SHAP-based explanations; and• explanation-rank stability assessment. Package contents The archive contains: • an executed notebook with saved analytical outputs;• a clean notebook for rerunning the workflow;• a Python script implementing the same analysis;• locked figures, tables, predictions, model outputs, calibration results, bootstrap estimates, sensitivity analyses, and explanation-stability outputs;• dataset-provenance and leakage-audit metadata;• model-hyperparameter, run-configuration, and software-environment manifests;• a validation report;• file-level SHA-256 checksums; and• environment specifications for reproducible execution. Data access and provenance Original input datasets are not redistributed in this archive. Source identifiers, retrieval guidance, retained-file characteristics, dimensions, licensing information, and SHA-256 checksums are documented in DATA_ACCESS.md and metadata/dataset_source_manifest_expanded.csv. The longitudinal analysis uses the subjective monitoring component of SoccerMon, an openly available dataset containing athlete-reported training load, wellness, injury, illness, and performance information. The relevant SoccerMon dataset record and its accompanying data-descriptor article are listed under Related works. Software version: 1.0.2 Licensing Original code and software components are distributed under the MIT License. SoccerMon-derived components retain attribution requirements associated with the Creative Commons Attribution 4.0 International license. Intended use This package is provided for research transparency, verification, and reproducibility. The resulting models are not presented as deployment-ready clinical decision-support systems and should not be used for individual medical or return-to-play decisions without independent external validation.

View free PDFSource page

Related papers

openalexZenodo (CERN European Organization for Nuclear Research)2026-07-23

Reproducibility Package for Explainable and Leakage-Conscious Machine Learning for Athlete Injury Risk Modeling Across Heterogeneous Datasets

Abdülkadir Enes GÖRGÜLÜ, Eray Dursun, Serdar SOLAK

This reproducibility package supports the manuscript “Explainable and Leakage-Conscious Machine Learning for Athlete Injury Risk Modeling Across Heterogeneous Datasets.” It contains the executed and clean analysis notebooks, the corresponding Python script, exact software-version…

View free PDFSource page
openalexZenodo (CERN European Organization for Nuclear Research)2026-07-24

Code and data for: Leakage-audited machine learning versus ETAS for earthquake forecasting in the Sea of Marmara

Basri Kerem Alhan, Kenessary Khabat

Code, processed data products, configuration, and results artifacts for "Machine learning versus ETAS for earthquake forecasting in the Sea of Marmara: a leakage-audited negative result and a closed-form scoring artifact" (Alhan & Khabat, submitted to Seismica). Version 1.2.0 acc…

View free PDFSource page
openalexZenodo (CERN European Organization for Nuclear Research)2026-08-09

A Systematic Review of Machine Learning, Deep Learning, and Explainable AI Approaches for Cardiac Disease Prediction

Sunanda Budihal, Sheetalrani Kawale, Abhishek Angadi

The cardiovascular (Cardiac) disease (CVD) is another factor that causes death among the global population most, and this is the reason why there is a high necessity to implement proper, effective, and interpretive diagnostic systems. The usage of machine learning (ML), deep lear…

Also available via: European Organization for Nuclear Research

View free PDFSource page
openalexZenodo (CERN European Organization for Nuclear Research)2026-07-23

Machine Learning Classification of COVID-19 Test Outcomes Using Symptom and Testing-Context Features

Edward Okhumaile

This preprint presents a reproducible machine-learning study of COVID-19 test outcome classification using symptom and testing-context features from a public 2020-2021 dataset. The study compares seven supervised learning algorithms and examines whether apparent model performance…

View free PDFSource page
openalexZenodo (CERN European Organization for Nuclear Research)2026-07-26

Comparative Analysis of Machine Learning Classification Algorithms and Hybrid Models for Student Performance Prediction

Ms. Pooja C. Soni, Dr. Hetal R. Modi, PC Negi

This study focuses on the analysis and comparison of machine learning classification algorithms and hybrid machine learning models for predicting student academic performance. Educational Data Mining techniques are used to extract meaningful insights from student datasets. Variou…

View free PDFSource page
openalexZenodo (CERN European Organization for Nuclear Research)2026-07-23

Village-Level Suitability Assessment For Litchi Cultivation In The Malwa Region Using Explainable Machine Learning And Geospatial Data

Dr. Pankaj Malik, Mishthi Patodia, Vedant soni, Deepika Kumari, Pragati Agrawal, Jaiswal Tanmay

Litchi is a high-value fruit crop traditionally cultivated in regions with favorable climatic and soil conditions. Expanding litchi cultivation into non-traditional areas such as the Malwa region of Madhya Pradesh requires accurate identification of suitable locations to minimize…

View free PDFSource page