律动BlockBeats
律动BlockBeats|7月 21, 2026 09:09
**[Kimi K3 Climbs to 4th on the Agent Leaderboard, Ranked First in User Confirmation Success Metric]** According to monitoring by 动察 Beating, the Arena Agent leaderboard is based on real user tasks and tool invocation records. Kimi K3 achieved a comprehensive net improvement of 9.62%, ranking 4th. It follows Claude Fable 5, Claude Opus 4.8 Thinking, and GPT-5.6 Sol. Arena evenly mixes all evaluated models into a virtual average benchmark and then estimates how much each metric improves when switched to K3. The final comprehensive score is the average of five net improvements. K3 has accumulated 8,344 test sessions. The user confirmation success metric saw a net improvement of 14.42%, ranking first. The "praise over complaints" metric improved by 20.62%, ranking third. Its error correction execution ranked 14th, and Bash error recovery ranked 17th. K3 is more likely to deliver results that users approve of, but mid-task error correction and command error recovery remain its weaknesses. [Original Link]
+5
Mentioned
Share To

Timeline

HotFlash

APP

X

Telegram

Facebook

Reddit

CopyLink

Hot Reads