Machine Learning-Based Prediction of CO2 Solubility in Deep Eutectic Solvents
Yiwen Wang, Shijia Peng, Chikezie Nwaoha, Teerawat Sema, Bin Liu, Paitoon Tontiwachwuthikul
Abstract Deep eutectic solvents, as an emerging class of green solvents, have demonstrated great potential in gas absorption and separation owing to their favorable physicochemical properties. However, accurate prediction of CO2 solubility in deep eutectic solvents across a wide range of temperatures and pressures remains a major challenge, limiting their optimization in carbon capture applications. In this work, two input representations, Simplified Molecular Input Line Entry System (SMILES)-based structural coding and physicochemical descriptors, were comparatively evaluated for CO2 solubility prediction in deep eutectic solvents. The dataset includes predominantly choline chloride-based deep eutectiv solvents, together with selected betaine-based and ammonium salt-based systems, spanning both hydrophilic and limited hydrophobic subclasses. The dataset contains 2,648 experimental measurements corresponding to 93 independent hydrogen bond acceptor-hydrogen bond donor (HBA-HBD) systems under different temperatures, pressures, and compositions. Four machine learning algorithms—extreme gradient boosting, random forest, deep neural network, and convolutional neural network—were evaluated using two input representations. All measurements associated with the same HBA–HBD pair were retained within the same data subset. Among the evaluated model–input combinations, the structural-coding-based random forest model achieved the highest test-set performance, with an R2 of 0.971 under the adopted random split. This study provides an exploratory comparison of machine learning strategies for CO2 solubility prediction within the deep eutectic solvent chemical space represented by the collected dataset.