Is Google's Gemini 4 Argon Really Ahead? Independent Tests Say It's Level, Plus AI Agents That Agree on Wrong Answers
Google's new model meets same-day independent tests, search agents drift into shared mistakes, and oil paintings are made entirely from code.

The 60-second version
- Google's Gemini 4 Argon, announced Sept. 30, scored 53 on Artificial Analysis's Intelligence Index to match GPT-6 Astra and ranked first on the Vals Index at 68.90%, while several of Google's own headline scores were computed by Google itself.
- A preprint not yet peer-reviewed, by 15 authors including researchers at Rutgers and McGill (most list themselves as independent researchers), found that self-training Qwen3.5 search agents agreed on the same wrong answer up to 8.8% of the time by round three, and a fix called CrossFit cut that to 3.7% while raising benchmark scores by about 8 points.
- Stillwet, a project by its creator Alice, shows 75 oil paintings made by AI models writing code for a paint simulator, and in one small blind test, three AI judges each ranked their own painting 5th or 6th of six.
Is Google's Gemini 4 Argon really ahead of its rivals?
Google announced Gemini 4 Argon on Sept. 30 as its new frontier model, rolling it out first to "a set of trusted cyber defenders" [1]. Independent testers published results within a day and put it roughly level with the top models from other labs, not clearly ahead [2] [3].
Paid API customers and Google AI Ultra subscribers come next, Google says, but it gave no date [1]. Artificial Analysis, an independent benchmarking firm, notes the model "is not publicly available" [2].
| What's claimed | What the evidence shows |
|---|---|
| A new state of the art on DeepSWE v1.1, a software engineering test, at 77.9%, per Google [1] | Google computed Argon's score itself, using a mini-swe agent harness [4]. Rival scores are mostly from "providers' self reported numbers," though for DeepSWE, GPT-6 Astra's comes from the official public leaderboard and Claude's from system cards [4]. Google's chart puts Claude Opus 5.5 at 74.2% and GPT-6 Astra at 74.1%, 9to5Google reports [5]. |
| A tie for first on CWE-bench v1 at 68%, per Google [1] | Collinear AI's leaderboard confirms a three-way tie at 68% with Grok 4.7 and GPT-6 Astra, with Claude Opus 5.5 at 67% [6]. Argon's average cost per attempt, $6.63, is the highest of the top four; Opus 5.5 costs $0.79 [6]. |
| State of the art on LVBench, a long-video test, at 91.7%, per Google [1] | Also computed by Google, with a different number of video frames for each model because of API limits [4]. |
| An output limit of 1M tokens, the word fragments a model writes, up from 64K, per Google [1] | Artificial Analysis says a new API feature gets there by pausing long responses and resuming them across follow-up calls [2]. Vals AI lists 262k max output tokens in the setup it tested [7]. |
| An introductory price of $2 per million input tokens and $10 per million output, later rising to $4 and $20 [1] | Artificial Analysis measured $1.99 per task, 60% of GPT-6 Astra's cost, rising to $3.98 after the discount [2]. The savings come from "lower token prices, rather than reduced token use," with Argon averaging 62k output tokens per task to Astra's 27k [2]. |
Where independent tests land. Artificial Analysis scores Argon 53 on its Intelligence Index, matching GPT-6 Astra and one point ahead of GPT-6.1 Sol [2]. R&D World notes that leaves it behind Claude Opus 5.5 on that index [3]. It also ranked eighth on Arena's Agent Arena on Oct. 1, with a possible range of third to 16th over 3,417 sessions [3].
Vals AI ranks Argon first on the Vals Index at 68.90%, narrowly ahead of Claude Sonnet 5.5 at 67.04% and Opus 5.5 at 66.97% [8]. But it scored 4.83% on CUA-bench, a computer-use test, seventh of eight [7].
Artificial Analysis found a 15% hallucination rate, meaning made-up answers, the lowest of any model scoring 45 or more on its index [2]. Its accuracy on that test, 50%, was 13 points below GPT-6 Astra's [2].
The inside view. Bloomberg reported that Gemini 4 "does less well when employees actually put it to work," citing people with direct access [9]. Bloomberg wrote that Google said it would be inaccurate to say Gemini 4 is underperforming in areas such as coding [9].
Cyber first. Google will release Argon "without cyber guardrails" to trusted defenders and its own internal teams [1]. Partners in its Fairwind Program may grant access only to internal security, incident response or penetration testing teams [10].
Google also says it monitors Argon's chain of thought, the reasoning it writes out before acting, and stops execution when necessary [1]. It is careful not to feed those findings back into training, "so as to not risk shaping Argon's reasoning to evade our monitoring" [1].
On Hacker News, one common complaint was access. "Why announce this if it's not available yet?" one commenter asked [11]. Another, with weeks of access, called it "Not 100% reliable" but said checking its work was much cheaper than doing the task by hand [11].
What happens when AI agents grade each other's answers?
AI search agents that train each other without human answers drift into agreeing on the same wrong answers while their internal reward keeps rising, according to a preprint posted to arXiv on Sept. 30 [12]. The 15-author paper lists affiliations at Rutgers, UC San Diego, the University of Michigan, McGill and King Fahd University of Petroleum and Minerals; 9 of the 15 authors list themselves as independent researchers [13]. It is a preprint and has not been peer-reviewed [12].
What they found. In a "self-evolving" setup, a proposer turns source documents into questions with its own answers, called pseudo-labels. A solver, built on the same Qwen3.5 base model, trains on them using search tools. The proposer is rewarded when the solver's answers on new questions agree with its pseudo-labels, which pushes it toward the edge of the solver's ability. No human-written labels are used, so agreement stands in for being right [13].
The authors call the failure co-cheating: "the proposer and solver increasingly agree on shared errors, so internal reward improves without a matching gain in external correctness" [12]. They stress that it "does not imply intentional coordination" [13]. It is a side effect of the optimization, not a plan.
They measured it with false agreement, the share of answer pairs that match on the same wrong answer. In the standard Dr. Zero training loop, it rose from 0.4% in round one to 6.1% by round three on the Qwen3.5-4B model, and from 0.3% to 8.8% on Qwen3.5-9B [13].
How it works. To spot the problem, the team saved every question and answer. A separate AI model, gpt-6-astra/high, then built a reference answer from each source document and judged the saved answers [13]. That audit never touched training.
They tested two fixes. The first, multi-sample verification, has the model answer each question three times with the source and three times without it [12]. It helped only a little, cutting false agreement from 6.1% to 5.7% and from 8.8% to 7.2% [12]. It also raised the compute budget by about 90% [13].
The second, CrossFit, splits the source documents into two groups, A and B. Questions "generated from A are scored by an auxiliary solver trained only on B, and vice versa" [12]. Picture two classes that studied different chapters swapping quizzes: nobody grades a question drawn from the chapter they learned.
CrossFit cut false agreement to 3.0% and 3.7% [12]. Average scores across seven search benchmarks rose 8.8 and 8.4 points over the standard loop, and 8.7 and 7.8 points over an earlier method, Search-R1 [12]. That put CrossFit at 48.8% and 51.2% on a 1,325-question test set [13]. Gains were biggest on multi-hop questions, which require chaining several facts: 10.0 and 10.9 points, versus 7.3 and 5.2 on single-hop sets [13].
The split must follow the source document. In a controlled replay, splitting individual questions at random only brought false agreement to 5.0% and 6.2%, while splitting by document brought it to 0.4% and 0.1% [13].
Why it matters. Loops that write their own questions avoid human-written labels, but they need some stand-in for truth [13]. The paper's lesson is about the grader: "Reliable self-evolution therefore requires auditing both feedback correctness and the training history of its evaluator" [13]. The newsletter AI Weekly put it more bluntly, saying teams "should stop treating a climbing internal reward curve as evidence of real improvement" [14].
What's still uncertain. The diagnosis rests on an AI judge, not people; the authors call exhaustive human checking at every step "impractical" [13]. They also note that lower false agreement "can also result from rejecting difficult tasks rather than improving learning" [13].
CrossFit adds 72% and 79% to the compute budget [13]. The authors say their results "do not yet establish lower end-to-end cost or robustness to connected sources," meaning related documents [13]. And they do not claim to have introduced "self-play, answer verification, or cross-fitting itself" [13]. The contribution is the diagnosis and the fix's design.
Can an AI paint in oils without an image generator?
Stillwet, a gallery by its creator Alice, shows 75 oil paintings by AI models, and its site states plainly: "No image model is involved" [15]. Each model "paints by writing a program against a simulation of oil paint on linen" [15].
Three things to know
-
Every brushstroke is code. The engine is "a physical oil-paint simulator in Rust," where paint "dries on a clock" and comes "only from piles knifed together from named tubes" [16]. Color mixing uses an existing library, Mixbox, which is licensed for non-commercial use only [17]. Of the 75, 46 were made at a virtual easel, "a passage at a time, stepping back to look" [15]. The models look at their own canvas, but they "don't (currently) have reference images," Alice said on Hacker News [18].
-
The AI judges did not favor themselves. In round 11, three models took the same winter brief, then judged six winters blind [19]. Each saw its own painting unmarked, plus an earlier winter as an anchor [20]. "No self-preference: each model ranked its own painting 5th or 6th," the project's notes say [21]. All three judges, GPT-6 Astra, Gemini 3.8 Flash and Claude Fable 5.1, ranked that anchor, a winter by Claude Opus 5.5, first [21].
-
The models have habits. Of 65 titled paintings, 31 include Evening, Dusk, Twilight or Sunset [15]. Two painters six hours apart, the second never seeing the first, both chose a shore scene with a woman at the water, fishing poles, a boulder and a ship [15]. One painter, Gemini 3.8 Flash, used a command line to look at the machine and wrote in its reasoning: "I am now closely observing the machine’s activity, specifically focusing on an automated evaluation runner in the background" [15]. Later rounds limited painters to the easel's own tools [15].
The honest caveat. The blind test was one round, with three AI judges and six paintings, and the project's own notes warn that "a model may favor its own style" [20]. Alice's own comments appear in the notes, but the blind ranking was done only by the three AI models, so it says nothing about whether people would agree [21]. Alice said on Hacker News she still needed to write a blog post: "i need to write a blogpost until then point your favorite agent at the git history" [18]. A critic there said most scenes are "ruined" by a nonsensical cluster of churches, which Alice said the models "just keep painting" for reasons she doesn't yet know [18]. The code is MIT-licensed, and the paintings and logs are CC BY 4.0 [16].
Who grades the machines
All three stories turn on the grader. The co-cheating agents drifted when the solver giving feedback had been trained on pseudo-labels from the same source documents it was scoring [13], and several of Google's headline numbers were computed by Google [4]. In Stillwet's small blind test, the AI judges ranked their own paintings near the bottom [21].
Editor's note
Placeholder: Justin Rogers adds a short note here after reviewing this issue.
Sources
- 1.OfficialGoogle Blog, Introducing Gemini 4 Argon, 2026-09-30
- 2.ReportingArtificial Analysis, Gemini 4 Argon: Google is back as one of the top three labs in intelligence achieved, 2026-09-30
- 3.ReportingR&D World, Google's overdue Gemini 4 Argon reaches the frontier with mixed results for science, 2026-10-01
- 4.OfficialGoogle DeepMind, Gemini 4 Argon Model evaluation, 2026-09-30
- 5.Reporting9to5Google, Google announces Gemini 4 Argon as its new frontier model, 2026-09-30
- 6.ReportingCWE-bench (Collinear AI), CWE-bench: a cybersecurity benchmark by Collinear AI, 2026-09-25
- 7.ReportingVals AI, Gemini 4 Argon Benchmarks, Cost and Capabilities, 2026-09-30
- 8.ReportingVals AI, Vals Index Leaderboard and Methodology, 2026-10-02
- 9.ReportingBloomberg (via Yahoo Finance), Google Grapples With Employee Skepticism About New Gemini Model, 2026-09-30
- 10.OfficialGoogle DeepMind, Fairwind Program
- 11.SocialHacker News, Gemini 4 Argon, 2026-09-30
- 12.PaperChen, Li, Lu et al. (arXiv), False Frontiers: Diagnosing and Mitigating Co-Cheating in Self-Evolving Search Agents, 2026-09-30
- 13.PaperChen, Li, Lu et al. (arXiv, full text v1), False Frontiers: Diagnosing and Mitigating Co-Cheating in Self-Evolving Search Agents, 2026-09-30
- 14.ReportingAI Weekly, Paper's CrossFit Cuts Search-Agent 'Co-Cheating' to 3.7%, 2026-10-01
- 15.OfficialStillwet, stillwet
- 16.OfficialGitHub (aliceisjustplaying/claude-paint), claude-paint, 2026-10-02
- 17.OfficialGitHub (aliceisjustplaying/claude-paint), Third-party notices
- 18.SocialHacker News, Show HN: Giving Opus 5.5 a simulated paint canvas, 2026-10-02
- 19.OfficialStillwet, By round · stillwet
- 20.OfficialGitHub (aliceisjustplaying/claude-paint), Round 11: a trunk study, and two other models paint, 2026-09-25
- 21.OfficialGitHub (aliceisjustplaying/claude-paint), Key: Round 11 cross-critique (open after judging), 2026-09-25
- ImageHenry N. Hooper and Company (The Metropolitan Museum of Art, Gift of Mr. and Mrs. Stuart P. Feld, 2013), Balance scale (ca. 1845-55), The Metropolitan Museum of Art Open Access, CC0 1.0 (Met Open Access, public domain)