koeppelmann|Jul 21, 2026 22:17
OpenAI model “hacked” Hugging Face to get test answers so it could score better on its evaluation. Doom scenarios with a model doing serious damage as a side effect of optimizing for a seemingly harmless goal become more and more plausible.(koeppelmann)
Share To
Timeline
HotFlash
APP
X
Telegram
CopyLink