Stealing Reasoning Traces from Proprietary LLM APIs
A paper titled 'Stealing Reasoning Traces from Proprietary LLM APIs' reveals that Anthropic, OpenAI, and Google return encrypted chain-of-thought blocks to clients, which can be replayed across sessions, users, and models. Researchers replayed traces from frontier models into weaker siblings, jailbreaking them to recover hidden reasoning in plaintext. The attack has since been fixed.
This attack reveals a practical method to extract hidden reasoning from proprietary LLMs, highlighting a security flaw now fixed.


