Evaluation awareness in language models: representation, verbalization, and control
Read the original at arxiv.org→arXiv:2608.21766v1 Announce Type: new Abstract: Both capability and safety benchmarks rest upon the assumption that the behavior of language models undergoing a test is informative about their behavior in...
Original headline: "Evaluation Awareness in Language Models: Representation, Verbalization, and Control"
Coverage timeline
- Aug 25, 04:00 UTC arXiv cs.CL lead source Evaluation Awareness in Language Models: Representation, Verbalization, and Control