Gemini 4 Argon Wins 12 of 18 Benchmarks. Code Isn't One.
DEV Community

Gemini 4 Argon Wins 12 of 18 Benchmarks. Code Isn't One.

Gemini 4 Argon Wins 12 of 18 Benchmarks. Code Isn't One. Google just dropped Gemini 4 Argon, and the numbers are hard to argue with. The model leads 12 of 18 published benchmarks against GPT-6 Astra and Claude Opus 5.5 -- scoring 77.9% on DeepSWE v1.1, 91.7% on LVBench video understanding, and landing the number one slot on LMArena Text Arena with 1,525 points. It ships a 1-million-token output window (up from 64K), hallucinates at a 15% rate where Astra hits 51%, and during its introductory period costs one-fifth what GPT-6 Astra charges. ๐Ÿ“– Read the full version with charts and embedded sources on ComputeLeap → But there is one column where Argon quietly underperforms: code. On Code Arena WebDev, it ranks 8th with 1,679 points -- trailing GPT-6.1 Sol (3rd, 1,759) by 80 points. On FrontierSWE v2, it scores 55.0% against Astra's 65.5%. On Terminal-bench 4.0, it finishes last among frontier models at 57.4%, nine points behind Claude Opus 5.5. That is not a bug. That is a strategy. The Benchmark Sweep: Where Argon Dominates The raw numbers paint an impressive picture. Here is how Argon stacks up against the other two frontier heavyweights across the benchmarks Google published: | Benchmark | Argon | GPT-6 Astra | Claude Opus 5.5 | Winner | |---|---|---|---|---| | DeepSWE v1.1 | 77.9% | 74.1% | 74.2% | Argon | | AutomationBench-AA | 77.5% | 41.4% | 42.5% | Argon | | GraphWalks | 84.2% | 71.8% | 66.8% | Argon | | LVBench (video) | 91.7% | 87.5% | 83.7% | Argon | | Vals Finance Agent v2 | 65.4% | 53.5% | 58.6% | Argon | | Harvey Legal Agent | 19.6% | 5.4% | 3.8% | Argon | | CWE-bench v1 | 68% | 68% | 67% | Tied | | FrontierSWE v2 | 55.0% | 65.5% | 62.1% | Astra | | Terminal-bench 4.0 | 57.4% | 62.1% | 66.4% | Opus 5.5 | | Code Arena WebDev | #8 (1679) | - | - | Sol #3 | The pattern is clear. On reasoning-heavy, knowledge-intensive, and multimodal benchmarks, Argon wins decisively. On AutomationBench -- which tests end-to-end business function execution -- it outscores both competitors by 35+ percentage points. On GraphWalks, it leads by 12-17 points. On long video understanding, it is the best model ever tested. โ„น๏ธ Argon leads the LMArena Text Arena at 1,525 points, 20 points ahead of Claude Opus 4.6 (High) in second. On the Artificial Analysis Intelligence Index, it scores 53 -- tied with GPT-6 Astra and Claude Fable 5.1, though Claude Opus 5.5 still leads at 58. But notice what is missing from that top cluster: from-scratch software engineering. View full article on VentureBeat → The Coding Gap Nobody at Google Is Talking About Code Arena WebDev is the benchmark that matters most to the developers actually reading this. It tests whether a model can build a real web application from a prompt -- not solve an isolated coding puzzle, but produce a working project. Argon lands in 8th place with 1,679 points, a 96-point improvement over Gemini 3.8 Flash (which sat at 29th), but still trailing the leaders by a meaningful margin. The gap shows up elsewhere too: - FrontierSWE v2: Argon scores 55.0% against Astra's 65.5% -- a 10.5 percentage point deficit on from-scratch project construction - Terminal-bench 4.0: Argon finishes last among frontier models at 57.4%, nine points behind Opus 5.5 (66.4%) - Terminal-bench Science 0.1: Argon hits 57.6% while Astra reaches 68.1% This is not a minor variance. On the benchmarks that matter most to working developers -- building software end-to-end -- Argon consistently trails by 9-10 points. โš ๏ธ Contrarian Corner: Google chose to publish 18 benchmarks, and Argon happens to lead on the 12 they selected. As AI researcher Horace He (@cHHillee) noted, "Everyone knows that labs can benchmaxx: optimize the model to look good on the leaderboard. Sometimes forgotten is that benchmarks can benchmaxx-maxx: optimize the benchmark to make its leaderboard look good." The benchmarks where Argon struggles -- the ones closest to what developers actually do -- are worth watching more closely than the ones it wins. The 1-Million-Token Output Window: Impressive, but Expensive Argon's headline feature beyond benchmarks is its industry-first 1-million-token output capability. Previous models capped at 64K output tokens. Argon can now produce responses 15x longer, using a "Long Decode Continuation" API that pauses and resumes across calls. The internal examples are legitimately impressive: processing 800K+ lines of the Fuchsia Zircon kernel, replacing 32,000 lines of SIMD code in the libgav1 video decoder (resulting in a 2.7x speed improvement), and identifying memory optimizations projected to free 300+ TiB across Google data centers. But there is a hidden cost. Analysis from The Decoder and Artificial Analysis reveals that Argon averages 62,000 output tokens per task compared to GPT-6 Astra's 27,000. That is 2.3x more tokens consumed per task. At introductory pricing ($2/$10 per 1M tokens), Argon costs $1.99 per Intelligence Index task versus Astra's $3.26 -- a genuine savings. But at regular pricing ($4/$20), that advantage narrows significantly. And the token verbosity raises a question: is Argon genuinely reasoning deeper, or is it just more verbose? The Pricing Play: Undercutting on Cost, Not Efficiency Google's pricing strategy is aggressive and deliberate: | Model | Input/1M | Output/1M | Cache Read/1M | |---|---|---|---| | Argon (intro) | $2 | $10 | $0.10 | | Argon (regular) | $4 | $20 | $0.20 | | Claude Opus 5.5 | $4 | $20 | $0.20 | | GPT-6 Astra | $10 | $50 | $1.00 | At promotional rates, Argon is one-fifth the price of Astra and half the price of Opus 5.5. The 95% cache read discount ($0.10 vs. $0.20 for Opus) makes it particularly attractive for workloads with heavy context reuse. But this is sticker-price competition, not efficiency competition. When you factor in Argon's 2.3x token consumption per task, the real-world cost comparison tightens. A developer running 1,000 tasks would pay roughly $1,990 with Argon versus $3,260 with Astra -- but at regular pricing, the gap closes to roughly $3,980 versus $3,260, actually making Astra cheaper per task. ๐Ÿ’ก For practitioners: if your workload involves long-context analysis, document processing, or agentic knowledge work, Argon's promotional pricing is a genuine bargain. But if you are running high-volume coding tasks, the per-task economics favor Claude or Astra due to Argon's higher token consumption. The Hallucination Tradeoff Nobody Mentions Argon's hallucination rate is outstanding: 15% on the AA-Omniscience benchmark, compared to GPT-6 Astra's 51%. That is 3.4x better. On the Gray Swan indirect prompt-injection test, attacks succeed against Argon just 0.7% of the time, versus 1.0% for Opus 5.5 and 8.5% for Astra. But Latent Space's analysis spotted an important tradeoff: Argon's accuracy on the same AA-Omniscience benchmark is 50%, compared to Astra's 63%. Argon hallucinates less -- but it also knows less. It is choosing to say "I don't know" more often, which improves the hallucination metric at the cost of overall accuracy. View full analysis on Latent Space → That is arguably the right engineering choice for enterprise deployments where a wrong answer is more dangerous than no answer. But it is worth understanding the mechanism: this is not a model that is both more accurate and less hallucinatory. It is a model that traded accuracy for reliability. What the Community Is Saying The Hacker News thread on Argon's announcement hit 1,592 points with 1,061 comments -- the largest AI model discussion on HN this quarter. The top comment captures the moment: "The important take away here: the leapfrogging we've seen this year doesn't seem to be a temporary thing. The famous theory of Dario Amodei was that AI was this winner-takes-all field where the first team to get a head start would never cede ground back." View full discussion on Hacker News → The community is split between genuine excitement about the benchmark numbers and skepticism about a model you cannot actually use yet. Hugging Face's Merve Noyan captured the mood of Google's supporters in two words: But one of the most upvoted HN comments points out the elephant in the room: "Gemini not beating the 'can't release a model' allegations." The prediction markets tell an interesting story too. Polymarket has Anthropic at 74% to have the best AI model at end of 2026, with Google at just 9.5%. The market is not fully buying the benchmark narrative without GA access. Karpathy's broader observation about LLM evaluation resonates here: "We're starting to leave the territory where you'd test an LLM by e.g. 'create an svg of pelican on a bicycle.'" The benchmarks are getting more sophisticated, but the gap between benchmark performance and real-world developer experience remains wide. Google's Strategic Bet: Win Everything Except Code Step back from the benchmark tables and Argon's positioning becomes clear. Google is making a deliberate choice about where to compete: Win: Reasoning, knowledge work, multimodal understanding, video, agentic business tasks, hallucination resistance, pricing. Concede: From-scratch software engineering, terminal-based coding, web app construction. This maps directly to what Sundar Pichai said on the Q2 2026 earnings call: Google is "a bit behind" in coding and agentic coding, and Gemini 4 was supposed to close that gap. Argon closes some of it -- the jump from 29th to 8th on Code Arena is real progress -- but the gap to the leaders remains substantial. The strategic logic is sound. Google's competitive advantage is in non-code AI applications: Search, Workspace, Cloud, Android, YouTube. Making Argon the best model for knowledge work, document understanding, and business automation serves those products. Code is where Anthropic (Claude Code, SWE-bench dominance) and OpenAI (Codex lineage, developer ecosystem) have structural advantages Google cannot easily replicate. โ„น๏ธ Google's internal validation supports this read. Argon's showcase results -- the Fuchsia kernel migra

Read on DEV Community ↗ ← Back to News

Comments

No comments yet. Start the discussion.