CORTEXA
← Browse
arxiveess.SY2026-06-27

Divergence-based Safety Measure for Large Language Models via Rational Inattention

Anh Tung Nguyen, Quanyan Zhu

This paper proposes a divergence-based safety measure for large language models (LLMs) under embedding-input attacks. The proposed measure quantifies the worst-case Kullback--Leibler divergence between the clean and attacked LLMs' output distributions, subject to a stealthiness constraint. This constraint is constructed by leveraging the equivalence between transformer attention used in LLMs and rational inattention modeling human decision-making. We analyze the proposed divergence-based safety measure by investigating perfectly undetectable attacks and deriving its upper bound through a Bregman-divergence argument. The proposed safety measure is applied to two pretrained causal language models, GPT-2 and GPT-Neo-125M, to show nontrivial output-distribution shifts, illustrating that the measure can distinguish model-level safety profiles.

View free PDFSource page

Related papers

arxivcs.ROeess.SY2026-07-01

Neuro-Symbolic Safety Guidance for Vision-Language-Action Models via Constrained Flow Matching

William English, Hao Zheng, Rickard Ewetz

Vision-Language-Action (VLA) models have demonstrated promising generalization capabilities across robotic manipulation tasks, yet their real-world deployment remains limited by the lack of effective safety measures. Specifically, existing safety measures only prevent collisions…

View free PDFSource page
arxiveess.SYcs.AI2026-06-30

Automating Cause-Effect Specification with Knowledge Graphs and Large Language Models

Javal Vyas, Milapji Singh Gill, Mehmet Mercangöz

Engineering specifications such as interlocks, alarm rationalization tables, and cause-and-effect (C&E) matrices remain central to process control and safety, yet their creation is still predominantly manual, document-driven, and prone to inconsistency. This paper presents a sema…

View free PDFSource page
arxivcs.ROcs.CVeess.SY2026-07-06

VLM-CASE: Vision-Language Model Enabled Context-Adaptive Safety Envelopes for Anticipatory Safe Autonomous Driving

Tianjia Yang, Ke Li, Ruwen Qin, Xianbiao Hu

Adverse driving conditions, such as bad weather, remain a principal barrier to autonomous driving because they degrade two things at once: what the vehicle can perceive and what it can physically do. Human drivers cope by anticipation, reasoning about the scene and re-budgeting s…

View free PDFSource page
arxiveess.SY2026-07-21

From P&ID Drawings to Process Graphs: A Multimodal Language Model Approach

Baikai Zhu, Samuel Duong, Javal Vyas, Mehmet Mercangöz

Piping and instrumentation diagrams (P&IDs) encode the functional structure of process plants and are a critical yet underutilised source of engineering knowledge for digital twins and intelli-gent decision support. However, digitising legacy P&IDs remains challenging due to hete…

View free PDFSource page