arxivcs.LG2026-07-01
On the Limits of Support-Preserving Alignment and Bounded Filtering
Aryan Dutt, Rui Mao, Anupam Chattopadhyay
We study whether alignment schemes that reshape a base model's output distribution, combined with bounded safety filters, can drive the probability of harmful behavior to zero in modern large language models. Recent research suggests that harmful behaviors can persist under prefe…