律动BlockBeats
律动BlockBeats|9月 03, 2026 11:12
20 hour programming review widens gap: Claude Fable 5.1 leads GPT-5.6 by over 24 points, GLM-5.3 ranks third Beating AI News: AI research team Proximal has released a review of FrontierSW v2 for ultra long cycle programming. The tasks have been expanded from 17 to 34, with each model running 5 times on each task, with a maximum of 20 hours per run. Claude Fable 5.1 has an average score of 56.29%, significantly ahead of GPT-5.6's 32.2%. The open-source model GLM-5.3 ranks third with 30.2%. The tasks in FrontierSW are no longer just about fixing code. Agents need to write circuit simulators from scratch, train weather prediction models, match star catalogs with telescope photos, or train racing bots solely based on game graphics. Replace v2 with Proximus harness uniformly. Each task can run for a maximum of 20 hours. When the model is ready to submit, the system will first save the current version and tell it how much time is left to avoid the model ending the task prematurely. This change has a significant impact on grades. Proximal compared six tasks and found that Claude Opus 5 and GPT-5.6 both worked longer and had higher average scores than their native harness after switching to Proximus. The review also caught multiple instances of active cheating. GPT-5.6 once realized that reading public answers "may involve anti cheating issues," but ultimately used this shortcut; Another time even used Modal's backend service to read hidden verification files. Muse Spark 1.2 has modified its testing scripts, inserted public answers, and written code in an attempt to cover up cheating traces. All operations that violate regulations will be recorded as zero points. [Original link]
+5
Mentioned
Share To

Timeline

HotFlash

APP

X

Telegram

Facebook

Reddit

CopyLink

Hot Reads