Deep Learning and Artificial Intelligence in the Design of Angiotensin-Converting Enzyme (ACE) Inhibitory Peptides
Zhangheng Qian, Yunuo Zhou, Ruotong Lin, Yang Hong, Yan Xu, S Li, 高江城, Qiurui Zhu, Heng Zheng
Background: Hypertension is a major modifiable risk factor for cardiovascular disease (CVD), and angiotensin-converting enzyme (ACE) remains a principal target within the renin-angiotensin-aldosterone system (RAAS). Food-derived ACE inhibitory peptides (ACEiPs) are attractive because of their specificity and generally favorable safety profile, but conventional discovery through protein hydrolysis, fractionation, purification, and repeated activity assays is labor-intensive and samples only a small fraction of the possible sequence space. Scope and approach: Unlike reviews that consider databases, predictive models, generative methods, or structural validation as separate topics, this review organizes these components into an integrated closed-loop workflow. It connects curated ACEiP data, sequence and chemical encodings, machine-learning and deep-learning predictors, de novo generation, structure prediction, docking, molecular dynamics, free-energy estimation, experimental validation, and feedback-guided model refinement. Key findings and conclusions: The analysis shows that database heterogeneity, inconsistent IC50 assay conditions, class imbalance, and sequence redundancy can limit model generalization. For short food-derived peptides, simple composition and physicochemical descriptors remain competitive baselines; protein language models (PLMs) can add contextual information but may be affected by protein-to-peptide domain shift, whereas SMILES and self-referencing embedded strings (SELFIES) are most useful when atom-level connectivity, stereochemistry, or non-canonical residues must be represented. A practical post-screening stage should jointly assess potency, ACE-domain binding, gastrointestinal stability, permeability, solubility, toxicity, allergenicity, and manufacturability. Progress toward translation therefore depends on a closed loop in which experimentally measured activity and developability data update the training set and guide subsequent design cycles.