Comparative Evaluation of Statistical, Machine‐Learning, and Deep‐Learning Models for Construction Sand and Gravel Price Forecasting: A Synthetic‐Data, Simulation‐Based Benchmark
You Wu, Peng Li, Jing Zhang, Fengsheng Guo, Danshu Hu, Wei He, Zhi Wang
ABSTRACT Forecasting construction sand and gravel prices is critical for infrastructure cost control, yet reliable comparisons among model families in the small‐sample, multidriver setting typical of regional markets are lacking. This study benchmarks eight algorithms—Ridge regression, Support Vector Regression (SVR), Gaussian Process Regression (GPR), Random Forest (RF), XGBoost, LightGBM, long short‐term memory (LSTM) networks, and an attention‐augmented LSTM (Attn‐LSTM)—on a fully synthetic 61‐month dataset. No observed Nanjing prices are used; the 15 domain‐relevant drivers are informed by market surveys but not calibrated to real data. The generating process is predominantly linear with mild Gaussian noise, so all rankings are conditional on this assumed mechanism. We use 6‐month lags, a 70/30 chronological split, and add classical baselines (constant mean, naive, seasonal‐naive, and an autoregressive integrated moving average [ARIMA; reference], reporting mean squared error [MSE], root mean squared error [RMSE], mean absolute error [MAE], percentage errors, and Diebold–Mariano tests). On the test set, SVR (MSE = 14.43, RMSE = 3.80) marginally beats Ridge (MSE = 15.58, RMSE = 3.95), but the difference is not statistically significant ( n = 17). Tree ensembles are mid‐tier, while LightGBM and GPR collapse to the training mean (MSE≈18.01). The two deep models, with ∼10 4 parameters for only 38 training sequences, produce the largest errors (MSE = 30.98 and 47.84). A univariate ARIMA (MSE = 12.78) outperforms all multivariate learners, underscoring that parsimony is rewarded here. We diagnose these failures. The main contribution is a transparent, reproducible simulation‐based benchmark and diagnostic comparison, not a validated real‐market forecasting tool.