AI and Machine Learning for Proteomics-Driven Drug Discovery: Methods, Tools, and Best Practices
Proteomics has become central to pharmacological research by providing quantitative readouts of protein abundance, post-translational modifications, interactions, and spatial context. However, proteomic datasets are high-dimensional, heterogeneous, and frequently affected by missingness, batch effects, and limited cohort size. Artificial intelligence (AI) and machine learning (ML) can help convert these complex data into decision-relevant outputs for target identification, biomarker discovery, pharmacodynamic monitoring, and drug repurposing. This review critically compares supervised learning, ensemble methods, dimensionality reduction, clustering, deep learning, graph learning, survival modeling, causal inference, and calibration approaches in proteomics-driven drug discovery. We also summarize major software ecosystems for mass-spectrometry processing, targeted assays, spectrum prediction, phosphoproteomics, structure modeling, and reproducible workflows. Emphasis is placed on model selection, benchmarking, missing-data handling, batch correction, interpretability, uncertainty, experimental validation, and translational readiness. Finally, we highlight emerging directions, including contrastive learning, diffusion models, graph-based integration, and federated analytics.