Tencent built its own Agent benchmark, yet its own CodeBuddy has never used Claude Code.
2026-08-11 17:22:51
According to CoinMeta, Tencent has built its own set of Agent benchmarks to test the performance of its own CodeBuddy and Claude Code. The 260 real-world problems are divided into four categories: coding, web, office work, and security. Each of the 7 models was tested using two sets of harness. The results showed that Claude Code won 17 out of the comparisons, while CodeBuddy only won 11. In the coding category, Claude Code emerged victorious with a score of 7:0. Although CodeBuddy had a slight 4:3 victory in the web and office work categories, it was defeated by Claude Code with a score of 4:3 in the security category. Tencent's ranking system also has its weaknesses, as there are only 50 to 80 questions in each category. The community has raised doubts about this, suggesting that adding more difficult questions could change the rankings.
Bullish 0
Bearish 0
Source:Internet
This content is for market information only and does not constitute investment advice.