Evaluating narrative unlearning with LENS: a level-based evaluation protocol for suppressing disinformation-aligned narrative reproduction in LLMs
Read the original at arxiv.org→arXiv:2607.22657v1 Announce Type: new Abstract: Large language models (LLMs) can reproduce disinformation-aligned narrative frames as plausible explanations, raising the question of whether existing...
Original headline: "Between Suppression and Collapse: Evaluating Narrative Unlearning with LENS"
Coverage timeline
- Jul 28, 04:00 UTC arXiv cs.CL lead source Between Suppression and Collapse: Evaluating Narrative Unlearning with LENS