CORTEXA
← Browse
crossrefAtmosphere2025-11-19Cited by 2

Integration of WRF-Chem Model-Based, Satellite-Based, and Ground-Based Observation Data to Predict PM2.5 Concentration by Machine Learning Approach

Soottida Chimla, Chakrit Chotamonsak, Tawee Chaipimonplin

Fine particulate matter (PM2.5) is a critical environmental and health concern in northern Thailand, where haze episodes are strongly influenced by biomass burning, meteorological variability, and complex topography. This study aims to (1) analyze and select input variables for PM2.5 prediction by integrating WRF-Chem outputs, satellite data, and ground observations, and (2) evaluate the predictive performance of four machine learning (ML) algorithms—Random Forest (RF), XGBoost, CNN3D, and ConvLSTM—during the 2024 haze season (January–May). The dataset included hourly PM2.5 observations from 54 stations, the WRF-Chem-simulated PM2.5 and meteorological variables, satellite-based fire data, and geographical data. To improve consistency with ground-based data, WRF-Chem PM2.5 values were bias-corrected for the training and validation phases prior to ML learning. Among Linear Regression, RF, XGBoost, Artificial Neural Network (ANN), and Convolutional Neural Network (CNN) tested for bias correction, RF achieved the best performance (R = 0.78, RMSE = 29.28 µg/m3); the RF-corrected WRF-Chem PM2.5 was then used as an input to the forecasting stage. Variable selection was supported by correlation, VIF, feature importance, and SHAP analyses. The results indicate that RF provided the most reliable predictions, achieving a correlation of R = 0.867 and the lowest RMSE of 27.6 µg/m3 when using the SHAP+VIF-selected input set (seven variables: PM2.5_lag1, PM2.5_lag24, T2, RH2, Precip, Burned Area, NDVI). Notably, RF remained the top performer, predicting PM2.5 more accurately than the other algorithms during high-pollution conditions, specifically Air Quality Index (AQI) “Unhealthy for Sensitive Groups” (high) and “Unhealthy” (very high). Taken together, RF set the performance bar across both stages, with XGBoost ranked second, whereas CNN3D and ConvLSTM performed considerably worse. These findings emphasize the effectiveness of ensemble tree-based algorithms combined with bias-corrected WRF-Chem outputs and strategic variable selection in supporting accurate hourly PM2.5 predictions for air quality management in biomass burning regions.

View free PDFSource page

Related papers

crossrefAtmosphere2024-09-29Cited by 12

Development of Machine Learning and Deep Learning Prediction Models for PM2.5 in Ho Chi Minh City, Vietnam

Phuc Hieu Nguyen, Nguyen Khoi Dao, Ly Sy Phu Nguyen

The application of machine learning and deep learning in air pollution management is becoming increasingly crucial, as these technologies enhance the accuracy of pollution prediction models, facilitating timely interventions and policy adjustments. They also facilitate the analys…

View free PDFSource page
crossrefAtmosphere2024-11-10Cited by 59

Systematic Review of Machine Learning and Deep Learning Techniques for Spatiotemporal Air Quality Prediction

Israel Edem Agbehadji, Ibidun Christiana Obagbuwa

Background: Although computational models are advancing air quality prediction, achieving the desired performance or accuracy of prediction remains a gap, which impacts the implementation of machine learning (ML) air quality prediction models. Several models have been employed an…

View free PDFSource page
crossrefAtmosphere2024-02-27Cited by 7

Analysis of Primary Air Pollutants’ Spatiotemporal Distributions Based on Satellite Imagery and Machine-Learning Techniques

Yanyu Li, Meng Zhang, Guodong Ma, Haoyuan Ren, Ende Yu

Accurate monitoring of air pollution is crucial to human health and the global environment. In this research, the various multispectral satellite data, including MODIS AOD/SR, Landsat 8 OLI, and Sentinel-2, together with the two most commonly used machine-learning models, viz. mu…

View free PDFSource page
crossrefAtmosphere2023-08-27Cited by 8

Downscaling Daily Satellite-Based Precipitation Estimates Using MODIS Cloud Optical and Microphysical Properties in Machine-Learning Models

Sergio Callaú Medrano, Frédéric Satgé, Jorge Molina-Carpio, Ramiro Pillco Zolá, Marie-Paule Bonnet

This study proposes a method for downscaling the spatial resolution of daily satellite-based precipitation estimates (SPEs) from 10 km to 1 km. The method deliberates a set of variables that have close relationships with daily precipitation events in a Random Forest (RF) regressi…

View free PDFSource page
crossrefAtmosphere2024-05-14

How Cloud Droplet Number Concentration Impacts Liquid Water Path and Precipitation in Marine Stratocumulus Clouds—A Satellite-Based Analysis Using Explainable Machine Learning

Lukas Zipfel, Hendrik Andersen, Daniel Peter Grosvenor, Jan Cermak

Aerosol–cloud–precipitation interactions (ACI) are a known major cause of uncertainties in simulations of the future climate. An improved understanding of the in-cloud processes accompanying ACI could help in advancing their implementation in global climate models. This is especi…

View free PDFSource page
crossrefAtmosphere2025-01-05Cited by 7

A Comparison of Machine Learning-Based Approaches in Estimating Surface PM2.5 Concentrations Focusing on Artificial Neural Networks and High Pollution Events

Shijin Wei, Kyle Shores, Yangyang Xu

Surface PM2.5 concentrations have significant implications for human health, necessitating accurate estimations. This study compares various machine learning models, including linear models, tree-based algorithms, and artificial neural networks (ANNs) for estimating PM2.5 concent…

View free PDFSource page