律动BlockBeats
律动BlockBeats|Sep 08, 2026 08:38
**[Meta Chief AI Officer Alexandr Wang Refutes SemiAnalysis' "Benchmaxxed" Claim: Old Tests Are Already Saturated]** Beating AI News Flash: AI research organization SemiAnalysis specifically called out Gemini 3.8 Flash and Muse Spark 1.3, labeling them as some of the most obvious "benchmaxxed" models—models overly optimized for public benchmarks. The two models scored 89.4% and 88.8% respectively on Terminal-Bench 2.1, approaching the performance of GPT-6 Astra and Claude Fable 5.1. However, on the newly released Terminal-Bench 4.0, their scores dropped drastically to 19.1% and 33.3%. In contrast, Astra and Fable 5.1 still managed 57.7% and 55.8%. SemiAnalysis suspects the issue lies in public benchmarks becoming increasingly susceptible to targeted optimization. All tasks in Terminal-Bench 2.1 are publicly available, meaning that even if model developers don't directly train on the test questions, they can purchase reinforcement learning datasets highly similar to these tasks. As a result, models become increasingly adept at excelling in these specific tests but may not generalize well to new tasks. However, SemiAnalysis has not provided direct evidence that companies like Google or Meta have purchased such datasets. Meta's Chief AI Officer Alexandr Wang subsequently refuted this claim, calling it "stupid." He cited an example where GPT-5.6 Sol also dropped from 88.8% on Terminal-Bench 2.1 to 37.3% on 4.0. His explanation is that 2.1 has already reached saturation and can no longer effectively differentiate top-tier models, whereas 4.0 is far from saturated. Terminal-Bench 4.0 indeed removed eight saturated tasks that had public solutions or other issues, modified 19 tasks, and readjusted constraints on time, CPU, and memory resources. [Original Link]
+5
Mentioned
Share To

Timeline

HotFlash

APP

X

Telegram

Facebook

Reddit

CopyLink

Hot Reads