A Unified Review of Statistical, Machine Learning, and Deep Learning Methods for Longitudinal Data Analysis
Oyebayo Ridwan Olaniran, Saheed Ajibade Kunle, Ali Rashash R. Alzahrani, Mohammed H. Alharbi, Nada MohammedSaeed Alharbi, Asma Ahmad Alzahrani
Longitudinal data, characterized by repeated measurements on the same subjects over time, are ubiquitous in biomedical sciences, economics, social sciences, and engineering. Analyzing such data presents unique statistical and computational challenges, including within-subject correlation, time-varying covariates, irregular observation times, informative dropout, and high dimensionality. While traditional statistical methods, such as linear mixed-effects models and generalized estimating equations, remain foundational, they often struggle with complex nonlinear dynamics, ultra-high-dimensional feature spaces, and very large sample sizes. Over the past two decades, machine learning (ML) and artificial intelligence (AI) methods have emerged as powerful complementary approaches to address these limitations. This review provides a comprehensive survey of mathematical and computational methods for longitudinal data analysis. We cover classical statistical models, penalized regression techniques, tree-based ensemble methods, kernel machines, Bayesian hierarchical models, and modern deep learning architectures, including recurrent neural networks, temporal convolutional networks, attention-based Transformers, neural ordinary differential equations, and generative models. We propose a unified taxonomy that organizes existing methods along two primary axes: the underlying mathematical framework and the analytical objective. For each category, we present detailed mathematical formulations, discuss key theoretical properties, examine computational considerations, and summarize representative reported applications drawn from the published literature. To increase the practical value of this review, we provide a cross-cutting comparison of method families against five key challenges (within-subject correlation, irregular sampling, missing data, high dimensionality, and scalability) and offer concrete guidance on method selection according to sample size, dimensionality, and analytical objective. Finally, we critically evaluate the strengths and limitations of these approaches, with particular emphasis on interpretability, scalability, handling of missing data, robustness to covariance misspecification, and uncertainty quantification.