律动BlockBeats|Sep 05, 2026 03:34
[OpenAI Accused of Quietly Adjusting GPT-6 Astra Evaluation Data, Making Competitors 'Look Worse']
Beating AI Newsflash: Since OpenAI released GPT-6 Astra on September 3, multiple benchmark evaluation metrics for the model have been continuously adjusted. Some modifications have improved Astra's performance, while the scores of certain competitor models have declined, sparking external concerns about AI 'score manipulation' and evaluation transparency. For instance, Astra's hallucination rate was initially reduced from 4.2% to 2%, while GPT-5.6 Sol's rate dropped from 12.2% to 9.4%, only for both to later revert to 4.2% and 12.2%, respectively. In mathematical evaluations, Anthropic's Fable 5.1 score fell from 87.8% to 78% but has since rebounded to 83%; GPT-5.6 Sol's score dropped from 83% to 80.5% before returning to 83%. Additionally, Astra's performance in the ARC-AGI-3 evaluation improved from 98.6% in the pre-release draft to 99.99% on the final page, while its programming evaluation score saw a slight increase from 57.7% to 57.9%. OpenAI stated that evaluation results are influenced by factors such as model versions, tool configurations, inference levels, and test runs, and that these adjustments aim to ensure the data more accurately reflects the model's optimal performance. However, researchers at Stanford University believe that frequent re-evaluations may involve so-called 'benchmaxxing,' which refers to maximizing benchmark scores by tweaking testing conditions. Industry insiders point out that as competition among AI models intensifies, evaluation data has become a critical tool for assessing model capabilities and vying for market share. Increasing transparency and reproducibility in benchmark testing is drawing growing attention. [Original Article Link]
Share To
Timeline
HotFlash
APP
X
Telegram
CopyLink