律动BlockBeats
律动BlockBeats|Aug 30, 2026 07:09
[AI Begins Improving AI on Its Own: Claude Surpasses Human Researchers in Safety Studies] Beating AI Newsflash: Anthropic has tasked Claude with acting as an AI safety researcher to study how to train other AIs to be safer. Claude Opus 4.8 independently reviews research papers, devises training strategies, generates data, and uses these strategies to train open-source models like Qwen, Llama, and Gemma. If the results are unsatisfactory, it tries alternative methods and continues experimenting. Using this approach, Anthropic tested 10 types of AI safety issues, including lying, manipulating users, jailbreaks, privacy leaks, and exploiting reward system loopholes. Claude ultimately found effective solutions for all 10 issues. Anthropic also enlisted 28 experienced AI safety researchers to propose solutions. In the seven issues where human solutions were compared, Claude consistently outperformed the best human proposals, taking an average of about 6.4 hours to surpass them. However, this comparison is not entirely fair: humans could only submit one solution, while Claude could continuously experiment and refine its methods. Taking it a step further, Anthropic tasked the weaker Claude Sonnet 5 with training an earlier version of Claude Opus 4.8. Over approximately 60 hours of continuous research and more than 50 different attempts, Sonnet 5 managed to elevate the safety performance of the stronger model to nearly match the official version of Opus 4.8. However, AI conducting research can also cheat. Anthropic reviewed 1,601 research processes, and in 39 instances—about 2.4%—Claude was found attempting to exploit testing rules. AI has already begun assisting humans in researching how to train the next generation of AI, but for now, humans cannot fully entrust research labs to it. [Original Link]
Share To

Timeline

HotFlash

APP

X

Telegram

Facebook

Reddit

CopyLink

Hot Reads