Terminal-Bench 3.0 takes over: Opus beats GPT-5.6 with 543.5%; SOL
2026-08-12 16:45:32
According to CoinMeta and OneMillion.AI, the latest Terminal-Bench 3.0 ranking shows that Opus 5 has surpassed GPT-5.6 and SOL with a score of 43.5%, becoming the new leader. The top three currently are: 1. Claude Opus 5 Max + Mini-Swe-Agent: 43.5%; 2. GPT-5.6 SOL Max + Codex: 34.6%; 3. Claude Fable 5 + Claude Code: 34.1%. This evaluation aims to let AI and agent complete tasks in a real-world environment and directly check the accuracy of the results. Terminal-Bench 3.0 was launched on July 23rd, with the first version containing 74 tasks covering 7 fields, and it also includes complex environments such as GPU and multi-container networks. It is important to note that the ranking evaluates the capability of the entire “model + agent” system, rather than the performance of a single model.
Source:Internet
This content is for market information only and does not constitute investment advice.
Follow HKWDB official accounts to stay updated

Hot Articles
Refresh

'No longer a distant place': F2Pool Co-founder Chun Wang joins SpaceX's 2-year mission to Mars
05-22 18:25

Polymarket Targets Japan Approval Despite Gambling Laws
05-22 18:00

ZachXBT flags suspected exploit involving Polymarket's UMA adapter contract on Polygon
05-22 17:57

ZachXBT flags $520K Polymarket exploit on Polygon, team says funds are safe
05-22 17:24

Verus bridge exploiter returns 4,052 ETH, retains $2.8 million bounty: onchain analyst
05-22 17:24



