CORTEXA
← Browse

Longwei Wang

2 papers indexed

arxivcs.CRcs.AI2026-07-08

Mechanistic Interpretability of LLM Jailbreaks via Internal Attribution Graphs

Anupam Wagle, Ifrat Ikhtear Uddin, Chaowei Zhang, Longwei Wang

Large language models (LLMs) exhibit remarkable capabilities but remain highly vulnerable to adversarial prompts and jailbreak attacks. Existing approaches primarily analyze these failures through input-output behaviors or attribution methods, offering limited insight into how ad…

View free PDFSource page
arxivcs.CVcs.AI2026-07-05

Explainable Novel Category Discovery in Semantic Concept Space

Ifrat Ikhtear Uddin, Yang Zhou, KC Santosh, Longwei Wang

Novel category discovery aims to identify unseen classes from unlabeled data by transferring knowledge from labeled categories, but most existing methods perform discovery in opaque latent feature spaces. As a result, they may separate novel categories accurately while providing…

View free PDFSource page