On the generalization of steering vectors for chain-of-thought faithfulness
Read the original at arxiv.org→arXiv:2607.29062v1 Announce Type: new Abstract: Model capabilities have improved in large part due to scaling chain of thought. This has been a promising development for AI safety--where models verbalize their...
Original headline: "On the Generalization of Steering Vectors for Chain-of-Thought Faithfulness"
Coverage timeline
- Aug 3, 04:00 UTC arXiv cs.AI lead source On the Generalization of Steering Vectors for Chain-of-Thought Faithfulness