DeflectBench benchmarks rhetorical fallacy generation in LLMs across deflection strategies and prompt framings
Read the original at arxiv.org→arXiv:2608.26119v1 Announce Type: new Abstract: Whether large language models can be prompted to generate rhetorical fallacies on demand, and whether current safety post-training constrains this behavior, has...
Original headline: "DeflectBench: A Benchmark for Evaluating Rhetorical Fallacy Generation in LLMs"
Coverage timeline
- Aug 28, 04:00 UTC arXiv cs.CL lead source DeflectBench: A Benchmark for Evaluating Rhetorical Fallacy Generation in LLMs