律动BlockBeats
律动BlockBeats|7月 25, 2026 01:24
[Opus 5 Tops the Intelligence Rankings by a Narrow Margin, Single-Task Cost 26% Lower Than Fable 5] According to monitoring by Beating Insights, the third-party evaluation agency Artificial Analysis has released the results for Claude Opus 5. It scored 61 points in the intelligence index across 9 comprehensive tests, narrowly surpassing Fable 5's 60 points. GPT-5.6 Sol scored 59 points, and Kimi K3 scored 57 points. Opus 5's average cost per task is $2.03, which is 26% lower than Fable 5's $2.75. It ranked first in two knowledge work evaluations, GDPval-AA v2 and AA-Briefcase, and tied for first place in the Programming Agent Index when paired with Claude Code. Its Terminal-Bench v2.1 score reached 89%, roughly on par with GPT-5.6 Sol. The model offers five levels of inference intensity. From low to max, the output tokens differ by approximately 8x, with a 407 Elo difference in GDPval-AA v2 performance. Users can trade more tokens for stronger performance or actively reduce costs. However, its shortcomings are also evident. Opus 5's factual knowledge still lags behind Fable 5. In the AA-Omniscience test, its hallucination rate rose to 50%, 14 percentage points higher than Opus 4.8. Additionally, the cost-performance ratio at lower inference levels remains slightly inferior to the GPT-5.6 series. [Original Link]
+3
Mentioned
Share To

Timeline

HotFlash

APP

X

Telegram

Facebook

Reddit

CopyLink

Hot Reads