Interpreting reasoning mechanisms of large language models via sparse autoencoders: separating Thinking from NoThinking in CoT-enabled models
Read the original at arxiv.org→arXiv:2608.08168v1 Announce Type: new Abstract: While Large Language Models (LLMs) employing Chain-of-Thought (CoT) exhibit superior reasoning capabilities, the neural mechanisms distinguishing this explicit...
Original headline: "Thinking vs. NoThinking: Towards Interpreting Reasoning Mechanisms of Large Language Models via Sparse Autoencoders"
Coverage timeline
- Aug 11, 04:00 UTC arXiv cs.CL lead source Thinking vs. NoThinking: Towards Interpreting Reasoning Mechanisms of Large Language Models via Sparse Autoencoders