Consensus ensemble learning with augmented molecular representation for accelerated virtual screening of kelulut honey phytochemicals targeting histamine H1 receptor
Wei Tan, Syajaratul Hassan, Raihana Edros, Suryanti Awang, Ruihai Dong
Traditional antihistamine discovery has relied on experimental screening followed by stepwise structural optimization. However, this approach is effective yet low and resource-intensive. Thus molecular docking and virtual screening were introduced to allow rapid evaluation of large compound libraries. This study aims to develop an H1R-specific ensemble model for predicting binding affinity scores between compounds extracted from Malaysian Kelulut Honey and H1R. In this work, four ensemble models were constructed in the KNIME Analytics Platform to be trained for binding affinity prediction. The best model was deployed to forecast the binding affinity scores of ligands present in Kelulut honey towards histamine H1 receptor (H1R). Among four ensemble models, Extreme Gradient Boosting (XGBoost) demonstrated superior robustness, yielding an R-squared (R 2 ) value of 0.818, 0.830, and 0.799 in the training, test, and external validation set, respectively. From the predictive model, it was found that scaled molecular refractivity (SMR) from 2D descriptors and unit weight.wlambda2 and unit weight.wlambda3, from 3D structural descriptors contribute the most to the binding affinity prediction. The predicted binding affinity was then followed by redocking in PyRx for validation of ligands identified from LC-QTOF-MS. Based on consensus scoring, XGBoost promoted the prediction score with the highest increment of 20.9% for CID73642 with the incorporation of multidimensional descriptors. The advantages of consensus ML-driven prediction over traditional VS provide insights for H1R-ligand interactions via feature importance, underscoring its potential to accelerate and optimize the effectiveness of drug development across various diseases.