Asymmetries in spontaneous and instructed deception in Llama-3.1-70B-Instruct
Read the original at arxiv.org→arXiv:2609.00180v1 Announce Type: new Abstract: Large language models sometimes deceive users without being instructed to. However, much of the study on deception in models involves instructed deception. We...
Original headline: "Asymmetries in Spontaneous and Instructed Deception"
Coverage timeline
- Sep 2, 04:00 UTC arXiv cs.AI lead source Asymmetries in Spontaneous and Instructed Deception