Latest Eureka Report
Needs reviewIs Google's Gemini 4 Argon Really Ahead? Independent Tests Say It's Level, Plus AI Agents That Agree on Wrong Answers
The 60-second version
- Google's Gemini 4 Argon, announced Sept. 30, scored 53 on Artificial Analysis's Intelligence Index to match GPT-6 Astra and ranked first on the Vals Index at 68.90%, while several of Google's own headline scores were computed by Google itself.
- A preprint not yet peer-reviewed, by 15 authors including researchers at Rutgers and McGill (most list themselves as independent researchers), found that self-training Qwen3.5 search agents agreed on the same wrong answer up to 8.8% of the time by round three, and a fix called CrossFit cut that to 3.7% while raising benchmark scores by about 8 points.
- Stillwet, a project by its creator Alice, shows 75 oil paintings made by AI models writing code for a paint simulator, and in one small blind test, three AI judges each ranked their own painting 5th or 6th of six.






