CORTEXA
← Browse
openalexZenodo (CERN European Organization for Nuclear Research)2026-07-24Cited by 0

Before the Model: Why Datasets and Data Representation Define What Machine Learning Can Learn

Jean Franck Loa Rojas

Machine learning systems do not learn reality directly; they learn from the representations preserved in their datasets. This structured narrative review examines how dataset purpose, coverage, integrity, labeling, independence, reproducibility, governance, and continuity determine what a model can legitimately learn and how its performance should be interpreted. The paper introduces an eight-dimensional dataset readiness framework designed as an operational decision layer before model training. Rather than replacing documentation frameworks such as Datasheets for Datasets, Data Statements, Dataset Nutrition Labels, or Data Cards, the proposed framework translates dispersed documentation and quality requirements into evidence-based readiness decisions: ready, conditionally ready, or not ready. A documentary application to the UCI Bank Marketing dataset demonstrates how readiness depends on the intended use. In particular, the post-call variable “duration” may be acceptable for retrospective benchmarking but constitutes target leakage when the objective is to prioritize customers before contact. The case illustrates why dataset readiness cannot be separated from the prediction moment, evaluation design, and deployment context. The framework is presented as a conceptual and practical contribution rather than a fully validated empirical instrument. Future work should evaluate inter-rater agreement, applicability across domains, and its effects on downstream model development and decision-making.

View free PDFSource page

Related papers

openalexZenodo (CERN European Organization for Nuclear Research)2026-07-23

Reproducibility Package for Explainable and Leakage-Conscious Machine Learning for Athlete Injury Risk Modeling Across Heterogeneous Datasets

Abdülkadir Enes GÖRGÜLÜ, Eray Dursun, Serdar SOLAK

This reproducibility package supports the manuscript “Explainable and Leakage-Conscious Machine Learning for Athlete Injury Risk Modeling Across Heterogeneous Datasets.” It contains the executed and clean analysis notebooks, the corresponding Python script, exact software-version…

View free PDFSource page
openalexZenodo (CERN European Organization for Nuclear Research)

Financial data source preparation and analysis as the initial stage of stochastic time series modeling by machine learning techniques

Yurchenko Yuriy, Oleksandr Zakovorotnyi

Proceedings of the Scientific Conference "The 13th International Scientific and Practical Online Conference of Young Scientists and Students ‘Contemporary Problems of Automation and Control’".The conference talk presents a structured approach to preparing and analyzing financial…

Also available via: European Organization for Nuclear Research

View free PDFSource page
openalexZenodo (CERN European Organization for Nuclear Research)

Feature Importance and Growth Rate Prediction in SiC PVT Processes through Advanced Machine Learning Models

Amir Reza Ansari Dezfoli

Silicon carbide is a key wide-bandgap semiconductor material for next-generation power electronics, yet the Physical Vapor Transport (PVT) method used for bulk crystal growth remains constrained by complex thermal-chemical interactions and low growth rates. This study develops a…

Also available via: European Organization for Nuclear Research

View free PDFSource page
openalexZenodo (CERN European Organization for Nuclear Research)

An Explainable Machine Learning Model and Bedside Nomogram Support Hemodialysis Decision-Making in Lithium Poisoning

Kamran Rezaei, Shahin Shadnia, Babak Mostafazadeh, Mitra Rahimi, Peyman erfantalabevini, Seyed Masoud Hosseini, et al.

The python codes for evaluation of a dataset containing lithium poisoning patients' data. Four Machine Learning models were used (Elastic-Net logistic regression (LR), linear support vector machine (SVM), shallow artificial neural network (ANN), and constrained Random Forest). Ea…

Also available via: European Organization for Nuclear Research

View free PDFSource page
openalexZenodo (CERN European Organization for Nuclear Research)2026-07-24

Simple Data Cleaning Techniques for Improving Machine Learning

Jyoti Panthangi

High-quality data is essential for building reliable machine learning models. Raw datasets often contain missing values, outliers, duplicates, inconsistent formats, and unstructured categorical variables. These issues reduce model accuracy and lead to biased predictions. This pap…

View free PDFSource page
openalexZenodo (CERN European Organization for Nuclear Research)2026-07-26

Enhancing Cardiovascular Disease Diagnosis through Data-Driven Feature Analysis and Cross-Validated Machine Learning Models

Abhilash Butola

Abstract - Cardiovascular diseases are a major global health problem, accounting for 17.9 million deaths per year and constituting 32 percent globally. According to the World Health Organization, the disease in people is due to an unhealthy diet,such as the intake of more junk fo…

View free PDFSource page