Wildfire Prediction in British Columbia Using Machine Learning and Deep Learning Models: A Data-Driven Framework
Maryam Nasourinia, Kalpdrum Passi
Wildfires pose a growing threat to ecosystems, infrastructure, and public safety, particularly in the province of British Columbia (BC), Canada. In recent years, the frequency, severity, and scale of wildfires in BC have increased significantly, largely due to climate change, human activity, and changing land use patterns. This study presents a comprehensive, data-driven approach to wildfire prediction, leveraging advanced machine learning (ML) and deep learning (DL) techniques. A high-resolution dataset was constructed by integrating five years of wildfire incident records from the Canadian Wildland Fire Information System (CWFIS) with ERA5 reanalysis climate data. The final dataset comprises more than 3.6 million spatiotemporal records and 148 environmental, meteorological, and geospatial features. Six feature selection techniques were evaluated, and five predictive models—Random Forest, XGBoost, LightGBM, CatBoost, and an RNN + LSTM—were trained and compared. The CatBoost model achieved the highest predictive performance with an accuracy of 93.4%, F1-score of 92.1%, and ROC-AUC of 0.94, while Random Forest achieved an accuracy of 92.6%. The study identifies key environmental variables, including surface temperature, humidity, wind speed, and soil moisture, as the most influential predictors of wildfire occurrence. These findings highlight the potential of data-driven AI frameworks to support early warning systems and enhance operational wildfire management in British Columbia.