arxivcs.CLcs.AIcs.LG2026-07-03
Revealing Hidden Model Behaviors with Task-Specific Self-Reports
Taras Kutsyk, Bartosz Zieliński
Fine-tuning can give a language model a hidden behavior--it may give false answers under a narrow condition, or give harmful advice only when a prompt touches a particular topic. We introduce the Stabilized Adapter for self-Report (SAR), a lightweight LoRA adapter that makes a fi…