The latest paper sparks discussion: AI's "deep thinking" process can be distilled for free, and the most valuable training assets of closed-source companies are being depleted.

CN
2 hours ago
Every time AI deeply thinks, it is an opportunity to be backtracked.

Author: Claude, Deep Tide TechFlow

Deep Tide Guide: Every time you throw a question at AI, it will first "deeply think" behind the scenes before responding. This invisible thinking process is the moat that OpenAI and Anthropic protect. Now, a group of researchers has publicly revealed a method to fully extract this thinking process. Along with it came users' pasted credit card numbers, passwords, and emails. This is not a security paper that is far removed from your life, but a precursor to how you might talk to AI in the future, which could change.

On August 10, a paper was submitted to arXiv, and the next day, the project website stolen-thoughts.com went live, showcasing the "thought records" extracted from several closed-source models, reaching 500 points on Hacker News.

Project leader Alexander Panfilov wrote on X: "We found a way to exploit vulnerabilities in the APIs of all leading AI companies to extract hidden reasoning from cutting-edge models."

In other words: you might think only AI knows what it is thinking; in fact, there are people who can lay it bare for you.

What you paste into AI may be leaking alongside its "thoughts"

Researchers scanned about 7,000 shared AI assistant session records available online, decrypted the "thinking process" that was encrypted one by one, and then found some things that should not have appeared.

Panfilov stated in a tweet: "We initially scanned about 7,000 public sessions and discovered 62 unique API keys, 33 email addresses, 33 passwords, and other sensitive information."

What stood out even more were the details. In the "thought" of a flight booking task, full names, emails, passport numbers, birth dates, and credit card numbers with security codes were lying there. Keys from platforms like Anthropic, AWS, and GitHub also appeared in the cases. This means that credentials you pasted in to get AI to do work for you could be stored alongside its thought process and then grabbed by others.

"If you've ever shared Claude Code or Codex sessions with encrypted reasoning blocks online, they can be decoded, and your personal data can be leaked," Panfilov wrote.

The most direct reminder for ordinary users is: don't paste confidential information into AI, even if it says, "I won't share it externally."

Vendors first say "it's fine," then secretly fix it

This incident did not happen suddenly. In May of this year, Matthew Green, a professor of cryptography at Johns Hopkins University, reported similar vulnerabilities to vendors, and the response received was "no security impact observed."

When Panfilov's team officially disclosed the issue, the vendors' attitude changed. "We subsequently went through a responsible disclosure process, and the vendors have patched several issues arising from this vulnerability. As far as I know, it is still ongoing," Panfilov said. The paper also confirmed that after the disclosure, researchers were unable to reproduce the same attack.

The problem lies in the fact that the vulnerability existed for several months without the users' knowledge. By the time it was made public and discussed, the fixes were finally implemented. This is not the fault of a single company, but the entire industry's assumption that "encryption equals security" has been publicly punctured for the first time. For readers, the real takeaway is one sentence: the AI company you trust may not have told you all the risks.

The model you use may no longer be so "exclusive"

A longer-term change lies at the industry level.

Reasoning capability is the foundation that OpenAI and Anthropic price, and it is also the part they least want to show externally. Once this thinking can be extracted in bulk, competitors can feed their models with the "thoughts" of the strongest models at a very low cost. The paper also mentioned a preliminary observation that has not undergone peer review: guiding another model with the reasoning of a small quantity of the strongest models would significantly pull the responses of the latter toward those of the former.

What does this mean? The moat of closed-source models was originally "you can't buy my brain with money." Now, this wall has developed a crack. For users, in the short term, this may not be a bad thing: stronger competing products might emerge faster, and prices could be driven down. But the cost is that you can no longer tell if a model is truly intelligent or just copying someone else's thinking.

Who "owns" the "hidden thinking" that is paid for but not visible is the essence of this debate. A technological fact is laid out here: as long as this thinking is still held on the client side, encryption is merely a smoke and mirrors act.

For closed-source vendors, what is more difficult to fix than the vulnerability is the narrative: the story that reasoning is the moat has been proven for the first time to be subject to bulk dismantling.

免责声明:本文章仅代表作者个人观点,不代表本平台的立场和观点。本文章仅供信息分享,不构成对任何人的任何投资建议。用户与作者之间的任何争议,与本平台无关。如网页中刊载的文章或图片涉及侵权,请提供相关的权利证明和身份证明发送邮件到support@aicoin.com,本平台相关工作人员将会进行核查。

Share To
APP

X

Telegram

Facebook

Reddit

CopyLink