As a Language Model...": Chat template switches LLM self-referential voice and activation steering reproduces it
Read the original at arxiv.org→arXiv:2609.25021v1 Announce Type: new Abstract: Large Language Models (LLMs) tend to add disclaimers like "I'm just an AI" when asked about something related to themselves. The self-reports from such responses are...
Original headline: ""As a Language Model...": Chat Template Switches LLM Self-Referential Voice and Activation Steering Reproduces It"
Coverage timeline
- Sep 23, 04:00 UTC arXiv cs.LG lead source "As a Language Model...": Chat Template Switches LLM Self-Referential Voice and Activation Steering Reproduces It