律动BlockBeats
律动BlockBeats|Aug 29, 2026 13:57
**[Terminal-Bench 4.0: GLM-5.3 Climbs to Third, Surpassing GPT-5.6 Sol]** Breaking AI News from Dongcha: Terminal-Bench has released version 4.0, recalibrating the time, CPU, and memory usage for Agent task execution, while fixing 19 tasks and removing 8 tasks due to saturation, refusal to answer, publicly available solutions, or quality issues. The maximum execution time for all tasks has been standardized to 8 hours, primarily to reduce the impact of timeouts and environmental issues on results. In the latest rankings, Opus 5 + Claude Code takes the top spot with 51.8%, followed by Fable 5 at 44.5%. GLM-5.3 + Claude Code achieved 41.8%, climbing to third place and surpassing GPT-5.6 Sol + Codex, which scored 37.3%. Among the top three, GLM-5.3 is the only model not developed by Anthropic. In Terminal-Bench 3.0, GLM-5.3 was ranked fourth with 32.4%, trailing GPT-5.6 Sol's 34.6%. By version 4.0, GLM-5.3 has risen to third place, now leading Sol by 4.5 percentage points. [Original Link]
+2
Mentioned
Share To

Timeline

HotFlash

APP

X

Telegram

Facebook

Reddit

CopyLink

Hot Reads