A comprehensive survey on feature representation for content based video retrieval
Abhijeet Jadhav, Sirshendu Arosh, Tamal Mondal, Prithviraj Pramanik
Content-based video retrieval (CBVR) has become an important research area due to the rapid growth in the video data. Such escalation took place due to the ubiquitous availability of internet access, IoT devices, smartphones, and cloud-based content sharing platforms like Snapchat, Instagram, YouTube, X, and many more. CBVR systems are extensively studied & utilized in past research that relied on feature representations to determine how video content is described, compared, and retrieved. The existing surveys are often restricted to specific time periods and application domains and focus on components like segmentation, keyframe extraction, dimensionality reduction, etc. However, none of the studies specifically highlights the evolution of feature representations across CBVR systems and applications. Thus, it appears to be challenging to obtain a temporal understanding of feature representations at various abstraction levels. The current study addresses this gap through introducing the evolution of feature extraction techniques, the datasets utilized, and the similarity measures that have been used in this evolution. Here, the study brings about the review of literature on CBVR systems published between 1995 and 2025. Considering the abstraction levels of feature classes, the literature is categorically discussed in terms of low, mid, and high-level feature representations. Further, the study explores different video datasets utilized by the literature and maps the influences of the various feature descriptors on the datasets. Finally, the study also traces the progress of feature representations adopted by different application domains. The findings reveal a clear progression in feature representation from handcrafted low-level descriptors to aggregated mid-level representations and, more recently, to deep learning-based high-level features, while also demonstrating the continued relevance of feature combinations and hybrid approaches. The significance of this review lies in its feature-centric and longitudinal perspective, which offers a single point of reference for comprehending earlier advancements in content-based video retrieval.