Before the Model: Why Datasets and Data Representation Define What Machine Learning Can Learn
Machine learning systems do not learn reality directly; they learn from the representations preserved in their datasets. This structured narrative review examines how dataset purpose, coverage, integrity, labeling, independence, reproducibility, governance, and continuity determine what a model can legitimately learn and how its performance should be interpreted. The paper introduces an eight-dimensional dataset readiness framework designed as an operational decision layer before model training. Rather than replacing documentation frameworks such as Datasheets for Datasets, Data Statements, Dataset Nutrition Labels, or Data Cards, the proposed framework translates dispersed documentation and quality requirements into evidence-based readiness decisions: ready, conditionally ready, or not ready. A documentary application to the UCI Bank Marketing dataset demonstrates how readiness depends on the intended use. In particular, the post-call variable “duration” may be acceptable for retrospective benchmarking but constitutes target leakage when the objective is to prioritize customers before contact. The case illustrates why dataset readiness cannot be separated from the prediction moment, evaluation design, and deployment context. The framework is presented as a conceptual and practical contribution rather than a fully validated empirical instrument. Future work should evaluate inter-rater agreement, applicability across domains, and its effects on downstream model development and decision-making.