CORTEXA
← Browse
crossrefMachine Learning and Knowledge Extraction2026-02-05Cited by 0

Enhancing the Extraction of GHG Emission-Reduction Targets from Sustainability Reports Using Vision Language Models

Lars Wilhelmi, Christian Bruns, Matthias Schumann

This study investigates how Vision Language Models (VLMs) can be used and methodically configured to extract Environmental, Social, and Governance (ESG) metrics from corporate sustainability reports, addressing the limitations of existing text-only and manual ESG data-extraction approaches. Using the Design Science Research Methodology, we developed an extraction artifact comprising a curated page-level dataset containing greenhouse gas (GHG) emission-reduction targets, an automated evaluation pipeline, model and text-preprocessing comparisons, and iterative prompt and few-shot refinement. Pages from oil and gas sustainability reports were processed directly by VLMs to preserve visual–textual structure, enabling a controlled comparison of text, image, and combined input modalities, with extraction quality assessed at page and attribute level using F1-scores. Among tested models, Mistral Small 3.2 demonstrated the most stable performance and was used to evaluate image, text, and combined modalities. Combined text + image modality performed best (F1 = 0.82), particularly on complex page layouts. The findings demonstrate how to effectively integrate visual and textual cues for ESG metric extraction with VLMs, though challenges remain for visually dense layouts and avoiding inference-based hallucinations.

View free PDFSource page

Related papers

crossrefMachine Learning and Knowledge Extraction2025-08-06Cited by 2

Evaluating Prompt Injection Attacks with LSTM-Based Generative Adversarial Networks: A Lightweight Alternative to Large Language Models

Sharaf Rashid, Edson Bollis, Lucas Pellicer, Darian Rabbani, Rafael Palacios, Aneesh Gupta, et al.

Generative Adversarial Networks (GANs) using Long Short-Term Memory (LSTM) provide a computationally cheaper approach for text generation compared to large language models (LLMs). The low hardware barrier of training GANs poses a threat because it means more bad actors may use th…

View free PDFSource page
crossrefMachine Learning and Knowledge Extraction2026-02-21

Plug-and-Play LLM Knowledge Extraction for Robot Navigation: A Fine-Tuning-Free Edge Framework

Sebastian Rojas-Ordoñez, Mikel Segura, Irune Yarza, Veronica Mendoza, Ekaitz Zulueta

Large Language Models are increasingly used for high-level robotic reasoning, yet their latency and stochasticity complicate their direct use in low-level control. Moreover, extracting actionable navigation cues from multimodal context incurs inference costs that are challenging…

View free PDFSource page
crossrefMachine Learning and Knowledge Extraction2026-01-04Cited by 1

Multivariate CO2 Emissions Forecasting Using Deep Neural Network Architectures

Eman AlShehri

One major factor influencing the development of eco-friendly policies and the implementation of climate change mitigation strategies is the accurate projection of CO2 emissions. Traditional statistical models face significant limitations in capturing complex nonlinear interaction…

View free PDFSource page
crossrefMachine Learning and Knowledge Extraction2024-02-16Cited by 14

VisFormers—Combining Vision and Transformers for Enhanced Complex Document Classification

Subhayu Dutta, Subhrangshu Adhikary, Ashutosh Dhar Dwivedi

Complex documents have text, figures, tables, and other elements. The classification of scanned copies of different categories of complex documents like memos, newspapers, letters, and more is essential for rapid digitization. However, this task is very challenging as most scanne…

View free PDFSource page
crossrefMachine Learning and Knowledge Extraction2025-07-16

Generalising Stock Detection in Retail Cabinets with Minimal Data Using a DenseNet and Vision Transformer Ensemble

Babak Rahi, Deniz Sagmanli, Felix Oppong, Direnc Pekaslan, Isaac Triguero

Generalising deep-learning models to perform well on unseen data domains with minimal retraining remains a significant challenge in computer vision. Even when the target task—such as quantifying the number of elements in an image—stays the same, data quality, shape, or form varia…

View free PDFSource page
crossrefMachine Learning and Knowledge Extraction2026-04-28Cited by 1

A Tiny Vision-Based Model for Real-Time Student Attention Detection in Online Classes

Chaymae Yahyati, Ismail Lamaakal, Yassine Maleh, Khalid El Makkaoui, Ibrahim Ouahbi

Online and blended classrooms widen access but remove the in-person cues instructors use to gauge attention. Prior work typically relies on heavy, cloud-bound or multimodal models that are hard to deploy on commodity laptops, treats attention as an unordered label without calibra…

View free PDFSource page