CORTEXA
← Browse
crossrefMachine Learning and Knowledge Extraction2024-02-16Cited by 14

VisFormers—Combining Vision and Transformers for Enhanced Complex Document Classification

Subhayu Dutta, Subhrangshu Adhikary, Ashutosh Dhar Dwivedi

Complex documents have text, figures, tables, and other elements. The classification of scanned copies of different categories of complex documents like memos, newspapers, letters, and more is essential for rapid digitization. However, this task is very challenging as most scanned complex documents look similar. This is because all documents have similar colors of the page and letters, similar textures for all papers, and very few contrasting features. Several attempts have been made in the state of the art to classify complex documents; however, only a few of these works have addressed the classification of complex documents with similar features, and among these, the performances could be more satisfactory. To overcome this, this paper presents a method to use an optical character reader to extract the texts. It proposes a multi-headed model to combine vision-based transfer learning and natural-language-based Transformers within the same network for simultaneous training for different inputs and optimizers in specific parts of the network. A subset of the Ryers Vision Lab Complex Document Information Processing dataset containing 16 different document classes was used to evaluate the performances. The proposed multi-headed VisFormers network classified the documents with up to 94.2% accuracy, while a regular natural-language-processing-based Transformer network achieved 83%, and vision-based VGG19 transfer learning could achieve only up to 90% accuracy. The model deployment can help sort the scanned copies of various documents into different categories.

View free PDFSource page

Related papers

crossrefMachine Learning and Knowledge Extraction2023-09-29Cited by 8

Optimal Topology of Vision Transformer for Real-Time Video Action Recognition in an End-To-End Cloud Solution

Saman Sarraf, Milton Kabia

This study introduces an optimal topology of vision transformers for real-time video action recognition in a cloud-based solution. Although model performance is a key criterion for real-time video analysis use cases, inference latency plays a more crucial role in adopting such te…

View free PDFSource page
crossrefMachine Learning and Knowledge Extraction2025-07-16

Generalising Stock Detection in Retail Cabinets with Minimal Data Using a DenseNet and Vision Transformer Ensemble

Babak Rahi, Deniz Sagmanli, Felix Oppong, Direnc Pekaslan, Isaac Triguero

Generalising deep-learning models to perform well on unseen data domains with minimal retraining remains a significant challenge in computer vision. Even when the target task—such as quantifying the number of elements in an image—stays the same, data quality, shape, or form varia…

View free PDFSource page
crossrefMachine Learning and Knowledge Extraction2025-10-01Cited by 4

Enhancing Cancer Classification from RNA Sequencing Data Using Deep Learning and Explainable AI

Haseeb Younis, Rosane Minghim

Cancer is one of the most deadly diseases, costing millions of lives and billions of USD every year. There are different ways to identify the biomarkers that can be used to detect cancer types and subtypes. RNA sequencing is steadily taking the lead as the method of choice due to…

View free PDFSource page
crossrefMachine Learning and Knowledge Extraction2025-07-16Cited by 4

Transformer-Driven Fault Detection in Self-Healing Networks: A Novel Attention-Based Framework for Adaptive Network Recovery

Parul Dubey, Pushkar Dubey, Pitshou N. Bokoro

Fault detection and remaining useful life (RUL) prediction are critical tasks in self-healing network (SHN) environments and industrial cyber–physical systems. These domains demand intelligent systems capable of handling dynamic, high-dimensional sensor data. However, existing op…

View free PDFSource page
crossrefMachine Learning and Knowledge Extraction2026-02-09

MERGE: Mammogram-Enhanced Representation via Wavelet-Guided CNNs for Computer-Aided Diagnosis of Breast Cancer

Omneya Attallah

The early and accurate identification of breast cancer is a significant healthcare issue, largely because the traditional machine learning approaches rely on handcrafted features that are unable to fully capture the spatial and textural complexity found in mammograms. Even with t…

View free PDFSource page
crossrefMachine Learning and Knowledge Extraction2025-12-01

SkinVisualNet: A Hybrid Deep Learning Approach Leveraging Explainable Models for Identifying Lyme Disease from Skin Rash Images

Amir Sohel, Rittik Chandra Das Turjy, Sarbajit Paul Bappy, Md Assaduzzaman, Ahmed Al Marouf, Jon George Rokne, et al.

Lyme disease, caused by the Borrelia burgdorferi bacterium and transmitted through black-legged (deer) tick bites, is becoming increasingly prevalent globally. According to data from the Lyme Disease Association, the number of cases has surged by more than 357% over the past 15 y…

View free PDFSource page