律动BlockBeats
律动BlockBeats|Sep 02, 2026 03:32
[Anthropic Locks Claude's Chain of Thought: Altering Context Invalidates It, Specifically Targeting Model Distillation] Beating AI News Flash: Anthropic has started adding 'context locks' to Claude's hidden thoughts. Encrypted thoughts generated by Fable 5.1 via API must now be returned exactly as they were, along with the system prompts, tools, and historical messages used during their generation. If any of the preceding content is altered, the API will either throw an error or discard the thought entirely. This change is aimed at preventing model distillation. Previously, researchers discovered that although Claude's encrypted thoughts were incomprehensible to users, they could be reinterpreted by Anthropic-compatible models. Attackers could first use Opus to generate high-quality reasoning, then feed the encrypted block to a less secure model like Haiku, tricking it into reproducing Opus's complete thought process. This essentially allowed them to not only copy the answers from the large model but also steal the step-by-step reasoning written on its 'scratch paper.' When Anthropic released Fable 5 in June, they had already implemented measures to bind encrypted thoughts to specific models, preventing them from being decrypted by smaller models like Haiku. With Fable 5.1, they have now added 'conversation binding': even if the model remains the same, any changes to the preceding prompts, tools, or chat history will render the old thoughts invalid. Previously, the focus was on preventing 'model-switching to steal thoughts,' but now even 'context-switching to reuse thoughts' has been blocked. [Original Article Link]
Share To

HotFlash

APP

X

Telegram

Facebook

Reddit

CopyLink

Hot Reads