Antropic, OpenAI, and Google return encrypted chain-of-thought blocks to clients that can be replayed across sessions and models; a trace from a frontier model can be replayed into a weaker model to jailbreak it and recover the stronger model’s hidden reasoning in plaintext.
Read the original at simonwillison.net→Stealing Reasoning Traces from Proprietary LLM APIs A vanity domain name (stolen-thoughts.com) for a neat paper: Anthropic, OpenAI, and Google return encrypted chain-of-thought blocks to clients that can be replayed...
Original headline: "Stealing Reasoning Traces from Proprietary LLM APIs"
Coverage timeline
- Aug 11, 22:40 UTC Simon Willison lead source Stealing Reasoning Traces from Proprietary LLM APIs
- Aug 11, 22:57 UTC Hacker News (AI) LLM Model-Swapping Trick Can Expose AI Reasoning Traces