律动BlockBeats
律动BlockBeats|8月 11, 2026 09:24
[Tencent Builds Its Own Agent Benchmark, Results Show CodeBuddy Losing to Claude Code] According to monitoring by Dongcha Beating, Tencent developed an Agent benchmark called WorkBuddy Bench and pitted its own CodeBuddy Code against Claude Code in a head-to-head comparison. A total of 260 real-world work tasks were divided into four categories: coding, web, office, and security. Seven models were tested using two different Harness setups. Harness refers to the agent execution layer outside the model, responsible for context management, tool invocation, and task execution. In 28 pairwise comparisons of the same model, Claude Code won 17 matchups, while CodeBuddy only won 11. In the coding category, Claude Code dominated with a clean sweep of 7:0—scores improved across all seven models when switched to Claude Code. However, CodeBuddy wasn’t completely outmatched. It achieved narrow victories in the web and office categories, both with a score of 4:3, while Claude Code took the security category with a 4:3 win. Interestingly, swapping only the Harness for the same model could result in score differences of over ten points. This indicates that evaluating the strength of an Agent can no longer rely solely on the underlying model. However, Tencent’s leaderboard has some clear weaknesses. After dividing the 260 tasks into four categories, each category only contained 50 to 80 tasks. Some community members directly questioned: Would the rankings change if more challenging tasks were added? Others on GitHub criticized the benchmark for testing only two Harness setups, noting that even Codex wasn’t included. [Original Link]
+2
Mentioned
Share To

Timeline

HotFlash

APP

X

Telegram

Facebook

Reddit

CopyLink

Hot Reads