Interpretable Machine Learning for Biomass Pyrolysis: Multi-Product Yield Prediction and Descriptor Prioritization for Bio-Oil, Syngas, and Biochar
Hira, Shagufta Zafar, Muhammad Imran Din, Farooq Ahmad, Wajahat Waheed Kazmi, Faysal M. Al-Khulaifi, Fiaz Hussain
Biomass pyrolysis produces three major product fractions—bio-oil, syngas, and biochar—whose yields are governed by nonlinear interactions among feedstock composition, catalyst properties, and operating conditions. Accurate prediction of these product yields remains challenging because literature-derived pyrolysis data are heterogeneous and unevenly distributed across the experimental space. In this study, an interpretable machine-learning framework was developed to predict biomass pyrolysis product yields using readily accessible descriptors from proximate/ultimate analysis, catalyst characterization, and process operation. A curated dataset containing 297 experimental records and 14 input descriptors was used to benchmark multiple regression algorithms, including linear and regularized linear models, support vector regression, k-nearest-neighbour regression, Random Forest, Extra Trees, gradient boosting regression trees, XGBoost, and CatBoost. Target-specific models were developed for bio-oil, syngas, and biochar to maximize the use of available yield data without imputing missing target values. The results showed that nonlinear tree-based ensembles outperform linear and distance-based approaches, while the optimal model depends on the product fraction. SHAP and feature-importance analyses further reveal distinct descriptor–yield relationships. Bio-oil prediction is mainly influenced by catalyst acidity, BET surface area, volatile matter, ash content, and temperature; syngas prediction is dominated by temperature and feedstock composition; and biochar prediction is strongly associated with heating rate, residence time, ash, fixed carbon, and oxygen content. This work provides a transparent data-driven tool for product-yield prediction and descriptor prioritization in biomass pyrolysis. The results showed that nonlinear tree-based ensembles consistently outperformed linear and distance-based approaches. On the independent test set, XGBoost achieved R2 = 0.834, RMSE = 5.503 wt%, and MAE = 4.371 wt% for bio-oil, and R2 = 0.790, RMSE = 5.427 wt%, and MAE = 3.791 wt% for syngas. GBRT achieved R2 = 0.739, RMSE = 3.966 wt%, and MAE = 2.706 wt% for biochar. Repeated five-fold cross-validation further produced mean R2 values of 0.783, 0.698, and 0.745 for bio-oil, syngas, and biochar, respectively.