CORTEXA
← Browse
openalexZenodo (CERN European Organization for Nuclear Research)2026-07-23Cited by 0

A causal perspective on Machine Learning for concrete quality predictions and data-driven mixture optimization

Thorsten Kalb, Anil Esen, Elsa Qoku, Thomas Matschei, Chiara Masiero, Gian Antonio Susto

Machine Learning (ML) predictions of cement and concrete quality and subsequent data-driven mixture optimization has been advertised for almost three decades. However, supervised ML leverages correlations, not causal relationships. Aiming for hybrid models, we derive the first causal mixture model for concrete, combining causal discovery and domain knowledge. This causal model highlights that mixture optimization is adversarial in cost, strength and workability; and omitting any of these yields trivial solutions. Challenging the literature consensus, we find that mixture design has therefore few degrees of freedom and strictly requires a causally interventional evaluation, rather than the correlation-based evaluation employed in most papers. Applying a perspective of formal causality, we identify common confounders and (spurious) correlations in different dataset types, among laboratory, literature, and production datasets. These considerations lead to specific guidelines for the experimental setup and evaluation of ML on concrete mixture data. A subsequent literature review reveals widespread malpractice and a quantitative link of apparently excellent model performance to poor evaluation practice. Finally, we demonstrate how these conceptual mistakes induce overestimation of model performance, optimal complexity, and usability by comparing the common spurious setup to our framework on two datasets. We find simple models to improve predictions of Compressive Strength in production, yet only moderately, whereas Slump is not predictable – contrary to the conclusions with the common setup. Overall, this work aims to advance data-driven mix design with ML by introducing causal reasoning to the field, raising awareness of ML limitations, and providing tools for unbiased performance evaluation.

View free PDFSource page

Related papers

openalexZenodo (CERN European Organization for Nuclear Research)2026-07-23

A Scalable Distributed and Fault-Tolerant Architecture for Cloud-Based Machine Learning and Data Analysis

Grace Dooshima GBOR, Emmanuel Ogala, Donald Douglas Atsa’am, Iorshashe Agaji

Abstract The rapid growth of data-intensive applications has necessitated the development of scalable and efficient architectures for cloud-based machine learning and data analysis. This study proposes a scalable, distributed, and fault-tolerant architecture designed to address t…

View free PDFSource page
openalexZenodo (CERN European Organization for Nuclear Research)2026-07-26

Enhancing Cardiovascular Disease Diagnosis through Data-Driven Feature Analysis and Cross-Validated Machine Learning Models

Abhilash Butola

Abstract - Cardiovascular diseases are a major global health problem, accounting for 17.9 million deaths per year and constituting 32 percent globally. According to the World Health Organization, the disease in people is due to an unhealthy diet,such as the intake of more junk fo…

View free PDFSource page
openalexZenodo (CERN European Organization for Nuclear Research)2026-07-24

Before the Model: Why Datasets and Data Representation Define What Machine Learning Can Learn

Jean Franck Loa Rojas

Machine learning systems do not learn reality directly; they learn from the representations preserved in their datasets. This structured narrative review examines how dataset purpose, coverage, integrity, labeling, independence, reproducibility, governance, and continuity determi…

View free PDFSource page
openalexZenodo (CERN European Organization for Nuclear Research)2026-07-24

Simple Data Cleaning Techniques for Improving Machine Learning

Jyoti Panthangi

High-quality data is essential for building reliable machine learning models. Raw datasets often contain missing values, outliers, duplicates, inconsistent formats, and unstructured categorical variables. These issues reduce model accuracy and lead to biased predictions. This pap…

View free PDFSource page
openalexZenodo (CERN European Organization for Nuclear Research)2026-07-23

A Machine-Learning-Driven Dataset of 140,000 PAH Infrared Spectra

Xinghong Mai

This ZIP archive contains the infrared (IR) spectral dataset of polycyclic aromatic hydrocarbons (PAHs) presented in the companion paper. The dataset comprises 144,111 IR spectra across 48,037 closed-shell, even-carbon benzenoid PAH structures in neutral, cationic, and anionic ch…

View free PDFSource page