Meta Chief AI Official Alexandr Wang Refutes SemiAnalysis "Ranking Brushing Theory": Old Tests Are Already Saturated
2026-09-08 16:23:28
According to CoinMeta, the chief AI official of Meta refuted the "rank-sweeping theory" put forward by the SemiAnalysis research institution, stating that the old tests had already reached saturation. SemiAnalysis pointed out that Gemini 3.8 and Muse Spark 1.3 are among the most prominent "benchmaxxed" models at present, achieving scores of 89.4% and 88.8% respectively in Terminal-bench 2.1, which is close to GPT-6 Astra and Claude Fable 5.1. However, in the newly released Terminal-bench 4.0, these two models only scored 19.1% and 33.3% respectively. Wang believes that Terminal-bench 2.1 has reached saturation and is no longer able to effectively distinguish between top models, whereas 4.0 has not yet reached saturation; moreover, some saturated tasks were removed and the testing conditions were adjusted.
Bullish 0
Bearish 0
Source:Internet
This content is for market information only and does not constitute investment advice.