On the limits of support-preserving alignment and bounded filtering; study examines whether alignment plus safety filters can drive harmful behavior to zero in large language models
Read the original at arxiv.org→arXiv:2607.18295v1 Announce Type: new Abstract: We study whether alignment schemes that reshape a base model's output distribution, combined with bounded safety filters, can drive the probability of harmful behavior...
Original headline: "On the Limits of Support-Preserving Alignment and Bounded Filtering"