25 models collectively lose to humans: AI still fails to understand many common life knowledge points
2026-10-08 19:33:43
According to CoinMeta and a report by OneMillion.AI, a new evaluation jointly conducted by research institutions scale, labs, and elorian under scale and ai shows that 25 multimodal models collectively lost to humans in the tests, being unable to understand many common aspects of life. The study covered 522 open-ended questions involving 288 images and 234 videos, with human participants achieving an accuracy rate of 93.1%. The top-ranked model gpt-6 astra only achieved an accuracy rate of 53.6%, while gpt-6.1 SOL and claude opus scored 46.6% and 44.6% respectively. The study found that 94% of the models' failures were related to missing key clues, misidentifying objects, or being unable to infer underlying relationships. In terms of social understanding, 21 models performed the worst, with video questions generally being more difficult than image questions. Despite increasing the amount of reasoning, the models still failed to match the performance of humans.
Bullish 0
Bearish 0
Source:Internet
This content is for market information only and does not constitute investment advice.