Lie typology, depth, and sparsity affect deception detection in LLM outputs
Read the original at arxiv.org→arXiv:2607.20479v1 Announce Type: new Abstract: Training probes to detect deceptive outputs from large language models is still an open problem. Recent work has demonstrated that detection probes fail especially in...
Original headline: "Beyond Liars' Bench: The Impact of Lie Typology, Depth, and Sparsity on Deception Detection in LLMs"
Coverage timeline
- Jul 24, 04:00 UTC arXiv cs.AI lead source Beyond Liars' Bench: The Impact of Lie Typology, Depth, and Sparsity on Deception Detection in LLMs