CORTEXA
← Browse
crossrefMachine Learning and Knowledge Extraction2023-09-29Cited by 8

Optimal Topology of Vision Transformer for Real-Time Video Action Recognition in an End-To-End Cloud Solution

Saman Sarraf, Milton Kabia

This study introduces an optimal topology of vision transformers for real-time video action recognition in a cloud-based solution. Although model performance is a key criterion for real-time video analysis use cases, inference latency plays a more crucial role in adopting such technology in real-world scenarios. Our objective is to reduce the inference latency of the solution while admissibly maintaining the vision transformer’s performance. Thus, we employed the optimal cloud components as the foundation of our machine learning pipeline and optimized the topology of vision transformers. We utilized UCF101, including more than one million action recognition video clips. The modeling pipeline consists of a preprocessing module to extract frames from video clips, training two-dimensional (2D) vision transformer models, and deep learning baselines. The pipeline also includes a postprocessing step to aggregate the frame-level predictions to generate the video-level predictions at inference. The results demonstrate that our optimal vision transformer model with an input dimension of 56 × 56 × 3 with eight attention heads produces an F1 score of 91.497% for the testing set. The optimized vision transformer reduces the inference latency by 40.70%, measured through a batch-processing approach, with a 55.63% faster training time than the baseline. Lastly, we developed an enhanced skip-frame approach to improve the inference latency by finding an optimal ratio of frames for prediction at inference, where we could further reduce the inference latency by 57.15%. This study reveals that the vision transformer model is highly optimizable for inference latency while maintaining the model performance.

View free PDFSource page

Related papers

crossrefMachine Learning and Knowledge Extraction2026-06-24

A Cognitive Lakehouse Framework with Transformer-Driven Analytics and Autonomous Decision Intelligence for Real-Time Enterprise Systems

Santosh Reddy Addula, Deepak Kumar, Guna Sekhar Sajja, Steven Hallman, Alan Dennis

The rapid evolution of data-driven enterprises demands scalable and intelligent systems capable of managing substantial volumes of heterogeneous data in real time. However, traditional systems lack a holistic approach to managing distributed data engineering, real-time analytics,…

View free PDFSource page
crossrefMachine Learning and Knowledge Extraction2026-04-28Cited by 1

A Tiny Vision-Based Model for Real-Time Student Attention Detection in Online Classes

Chaymae Yahyati, Ismail Lamaakal, Yassine Maleh, Khalid El Makkaoui, Ibrahim Ouahbi

Online and blended classrooms widen access but remove the in-person cues instructors use to gauge attention. Prior work typically relies on heavy, cloud-bound or multimodal models that are hard to deploy on commodity laptops, treats attention as an unordered label without calibra…

View free PDFSource page
crossrefMachine Learning and Knowledge Extraction2026-05-11

Knowledge Graphs in Autonomous Driving: Construction, Integration, and Real-Time Reasoning

Patrik Viktor, Gábor Kiss

Autonomous driving systems require the integration of heterogeneous sensor data, distributed V2X communication, and safety-critical decision-making into coherent and interpretable world models. This review provides a systematic analysis of knowledge graph (KG)-based approaches in…

View free PDFSource page
crossrefMachine Learning and Knowledge Extraction2026-03-09Cited by 2

Drift-Aware Online Ensemble Learning for Real-Time Cybersecurity in Internet of Medical Things Networks

Fazliddin Makhmudov, Gayrat Juraev, Ozod Yusupov, Parvina Nasriddinova, Dusmurod Kilichev

The rapid growth of Internet of Medical Things (IoMT) devices has revolutionized diagnostics and patient care within smart healthcare networks. However, this progress has also expanded the attack surface due to the heterogeneity and interconnectivity of medical devices. To overco…

View free PDFSource page
crossrefMachine Learning and Knowledge Extraction2026-04-15Cited by 1

Lightweight Deep Learning Models for Face Mask Detection in Real-Time Edge Environments: A Review and Future Research Directions

Saim Rasheed

Automated face mask detection remains an important component of hygiene compliance, occupational safety, and public health monitoring, even in post-pandemic environments where real-time and non-intrusive surveillance is required. Traditional deep learning models provide strong re…

View free PDFSource page
crossrefMachine Learning and Knowledge Extraction2024-01-26Cited by 18

Real-Time Droplet Detection for Agricultural Spraying Systems: A Deep Learning Approach

Nhut Huynh, Kim-Doang Nguyen

Nozzles are ubiquitous in agriculture: they are used to spray and apply nutrients and pesticides to crops. The properties of droplets sprayed from nozzles are vital factors that determine the effectiveness of the spray. Droplet size and other characteristics affect spray retentio…

View free PDFSource page