Committed before reasoning: evidence of answer pre-commitment in an open-weight LLM with behavioral reproduction and preliminary activation-level findings
Read the original at arxiv.org→arXiv:2607.16451v1 Announce Type: new Abstract: Chat models sometimes commit to an answer and then produce reasoning that justifies it rather than deriving it -- even when the answer contradicts a task premise. We...
Original headline: "Committed Before Reasoning: Behavioral Reproduction and Preliminary Activation-Level Evidence of Answer Pre-Commitment in an Open-Weight LLM"