arxivcs.CRcs.AI2026-07-14
Silent Alarm: A J-Space Protocol for Comparing Danger Recognition Across Models and Quantization Levels
Roman Prosvirnin, Victor Minchenkov, Alexey Soldatov, Vladimir Bashun
Jailbreak-robustness research typically evaluates safety through generated responses using an LLM-as-judge approach. Such evaluations, however, are sensitive to the benchmark's grading procedure and capture only observed behavior on a given set of attacks, without directly reveal…