CleaveSmart: deciphering the puzzle of rational 10–23 DNAzyme selection through interpretable AI insights
Fatemeh Rahimpour, Fatemeh Javadi‐Zarnaghi, Elahe Mousavi, Vajihe Akbari
The 10–23 DNAzyme is a versatile platform for programmable RNA cleavage in gene regulation and diagnostics; however, its practical efficacy remains limited by the absence of precise selection criteria for optimal enzyme–substrate pairs. While machine learning (ML) offers a robust framework for addressing this complexity, previous computational approaches have lacked sufficient predictive power and generalizability. Here, we present a feature-centric ML framework developed using a comprehensive dataset of 38 DNAzyme–substrate interactions. By engineering an expanded suite of physicochemical descriptors—integrating binding energy, DNAzyme internal and homodimer energies, and RNA cost energy—we evaluated multiple algorithms, with the Random Forest model demonstrating superior performance. Our optimized model accurately predicts sustained catalytic activity (60-minute cleavage), as validated on an independent experimental dataset of 11 samples (Accuracy = 0.91, F1-score = 0.95). Notably, a significant decline in predictive power for rapid activity (10-minute) suggests a mechanistic divergence: while sustained catalysis is primarily governed by thermodynamic stability, early-phase cleavage is likely dominated by transient kinetic factors. We integrated these insights into CleaveSmart (cleavesmart.streamlit.app). This transparent, open-access pipeline replaces empirical trial-and-error with data-driven selection, significantly facilitating the identification of high-potency DNAzymes.