OpenAI accused of quietly adjusting GPT-6Astra evaluation data, making some indicators 'worse' for competitors
2026-09-05 11:38:47
According to CoinMeta, OpenAI is alleged to have quietly adjusted the evaluation data for GPT-6 and Astra. Since the release on September 3rd, multiple model evaluation benchmark data points have been altered, with some modifications making Astra perform better while the scores of competing models declined, raising doubts about the transparency of AI's evaluations. The illusion rate of Astra once dropped from 4.2% to 2%, and that of GPT-5.6 and SOL decreased from 12.2% to 9.4% before returning to their original values. In mathematical evaluations, the Anthropic's Fable score of 5.1 fell from 87.8% to 78% but has now rebounded to 83%. Additionally, Astra's score in the ARC-AGI-3 evaluation increased from 98.6% to 99.99%. OpenAI stated that the evaluation results are affected by factors such as model versions and tool configurations, and this adjustment was aimed at ensuring that the data accurately reflects the model's performance. Researchers from Stanford University believe that frequent re-evaluations may involve "cheating," which is done by adjusting test conditions to maximize benchmark scores. As competition among AI models intensifies, evaluation data has become an important tool for measuring model capabilities, and transparency and reproducibility have come under scrutiny.
Bullish 0
Bearish 0
Source:Internet
This content is for market information only and does not constitute investment advice.