Prediction and Interpretation of the Volumetric Mass Transfer Coefficient in Bioreactors Using a No-Code Platform for Autonomous Machine Learning Model Selection
Ho-Yeon Lee, Yonghee Shin, Jongsun Won, Jin Ho Lee, Sangmin Park, Sang-Min Paik, Hwa Sung Shin, Moo Sun Hong, Jun-Woo Kim
The volumetric mass transfer coefficient (kLa) governs the design, operation, and scale-up of aerobic bioprocesses, yet its dependence on reactor geometry, impeller design, operating conditions, and fluid properties limits prediction by empirical correlations. Machine learning (ML) improves accuracy but faces two barriers in bioprocess practice: selecting the best model among many candidates requires expertise, and small, highly multicollinear data make models chosen based on test error alone prone to overfitting. Using a browser-based, no-code platform, we trained 14 regression algorithms under an identical pipeline on a published kLa dataset, and introduced a composite objective, the generalization-penalized error (GPE), which is the test RMSE plus the absolute train–test RMSE gap. Minimizing GPE rather than test RMSE expanded the top statistically equivalent group to include not only boosting ensembles but also simpler, interpretable models, indicating that black-box models hold no clear advantage once train–test consistency is assessed. Sensitivity analysis showed that tree models produce discontinuous responses, whereas algebraic learning via elastic net (ALVEN) yields smooth surfaces. Shapley additive explanations (SHAP) and an ontology graph, interpreted by a retrieval-augmented language-model agent, identified rotational speed and gas flow rate as dominant, reproducing the established mass transfer mechanism. The framework offers a reproducible, interpretable, expertise-light route to bioprocess model selection.