深潮TechFlow|Sep 05, 2026 03:29
[OpenAI Accused of Adjusting GPT-6 Astra Evaluation Data, Significant Changes in Some Metrics Reported]
According to Deep Tide TechFlow on September 5, as reported by *Fortune*, OpenAI has officially released GPT-6 Astra, calling it "the world's most intelligent and highly aligned model to date." However, the company has repeatedly modified the evaluation data on the model's performance page, with significant changes observed in some metrics.
Reportedly, Astra's internal "hallucination rate" was initially announced as 4.2%, later adjusted to 2%, and then reverted back to 4.2%. Meanwhile, Anthropic's Fable 5.1 score in FrontierMath Tier 4 was reduced from 87.8% to 78% and has since been restored to 83%. Additionally, Astra's performance on ARC-AGI-3 improved from the pre-release score of 98.6% to the current page's displayed score of 99.99%.
OpenAI responded by stating that the company places a high priority on evaluation accuracy, and variations of a few percentage points can result from differences in model checkpoints, tool configurations, inference levels, and testing methods. The adjustments aim to ensure that the evaluation data more accurately reflects the model's actual performance.
Market analysts believe that the AI industry has long been plagued by a phenomenon known as "benchmaxxing," where testing conditions and methods are repeatedly adjusted to maximize benchmark test scores.
Share To
Timeline
HotFlash
APP
X
Telegram
CopyLink