AI Weekly: Four Frontier Models in Four Days
Week of August 11 to 18, 2026 Four labs shipped frontier models within four days of each other this week, and every one of them was tuned for the same thing: agents that stay on task. SpaceXAI released Grok 4.6 and closed its Cursor acquisition, Google shipped Gemini 3.7 Flash at half price, DeepSeek took V4 Pro to general availability and then raised its prices, and Z.ai announced GLM-5.3 with cybersecurity claims that real CVE databases partially back up. Below the model layer, the MCP stateless spec entered its adoption window, and the memory market quietly delivered the most consequential news of all: 2027 DRAM and HBM capacity is reportedly already sold out. As always, the order is models first, then tooling, then standards, then infrastructure. Models set what is possible, tooling determines who can use it, standards decide whether the pieces connect, and infrastructure sets the cost. Models: Grok 4.6, Gemini 3.7 Flash, DeepSeek V4 Pro GA, and GLM-5.3 Grok 4.6 bets everything on long-horizon agents SpaceXAI released Grok 4.6 on August 12, and the release notes read like a thesis statement about where frontier labs think the value is. This is a post-training upgrade over Grok 4.5 rather than a larger base model. The lab held the foundation constant and spent the improvement budget on a longer supplemental training run, regenerated supervised fine-tuning trajectories, and reinforcement learning inside agentic environments. The goal is agents that stay on a task across many steps without drifting. The specs: a 500,000-token context window, a new xhigh reasoning-effort level above the existing ladder, and tiered pricing at $2 per million input tokens, $0.50 for cached input, and $6 per million output tokens below 200K prompt tokens. Above that threshold, prices double to $4, $1, and $12. The model is generally available through the xAI API as grok-4.6, is the default model in Grok Build, and ships in Cursor with doubled included usage for the first week. The independent numbers are genuinely interesting. Artificial Analysis scores Grok 4.6 at 61 on its Intelligence Index, up five points from Grok 4.5 and tied with GPT-5.6 Sol Max for third place overall. On AA-Briefcase, a long-horizon professional work benchmark, it posts an Elo of 1,577, narrowly above Claude Fable 5 Max at 1,574. The efficiency story stands out even more: Artificial Analysis reports Grok 4.6 completed its AA-Briefcase workloads in roughly 53 turns and about 0.5 billion input tokens on average, against roughly 103 turns and 2 billion input tokens for Claude Opus 5 Max. Fewer turns means less re-read context on every step, which compounds into real cost savings for production agents. Now the honest caveats. The bolded wins on GDPval-AA v2 and AA-Briefcase sit inside published confidence intervals, so they are statistical ties rather than leads. The comparison set in SpaceXAI's own table excludes Claude Opus 5, which currently tops the Artificial Analysis index at 63. And on the coding rows engineering teams care about most, Grok 4.6 still trails: 65.9% on DeepSWE v1.1 against 73% for GPT-5.6 Sol Max, and 26% on Terminal-Bench v3.0, nearly double its predecessor and still last among the listed frontier models. Artificial Analysis also places it at $0.84 per completed task, less economical than GPT-5.6 Luna and GLM-5.2. Grok 4.6 is a real step forward for long-running agent work and an incomplete one for coding. Gemini 3.7 Flash: coding gains at half price, for now Google released Gemini 3.7 Flash on August 13, 23 days after Gemini 3.6 Flash, and priced it to move. The introductory rate is $0.75 per million input tokens and $3.75 per million output tokens, with output charges including thinking tokens. That pricing expires on December 31, 2026, after which the rate doubles to $1.50 and $7.50, exactly what 3.6 Flash cost at launch. Google also applied the promotional rate to 3.6 Flash, so through year-end the migration decision is about capability, not list price. The specs are unchanged from 3.6 Flash: a 1,048,576-token input context window, a 65,536-token output limit, a March 2026 knowledge cutoff, and multimodal input across text, image, video, audio, and PDF with text output. The API exposes tunable thinking levels of low, medium, and high, and returns an error on the unsupported minimal setting. Availability spans the Gemini API, Google AI Studio, Antigravity, Android Studio, Gemini Enterprise, and Gemini Spark. The benchmark story is all software engineering, and every headline number Google published is a coding or automation test. The flagship result is DeepSWE v1.1 at 65.3%, against 49.0% for Gemini 3.6 Flash, a 16-point generational jump on a long-horizon software engineering eval. Google-reported numbers also show GDM-MRCR v2 long-context retrieval improving from 91.8% to 97.0% at 128K, and OSWorld-2.0 computer use rising from 33.8% to 47.9%. Those figures are vendor-reported, so treat them as release evidence rather than independent results. On the independent side, Artificial Analysis scores the model 56 on its Intelligence Index against 52 for 3.6 Flash, and ranks it first of 186 models on output speed at 340.1 tokens per second. GPT-5.6 Terra still leads on DeepSWE, Terminal-Bench, and OSWorld in cross-vendor comparisons. The practitioner takeaway: a model that resolves an agentic task in fewer intermediate steps saves both the output tokens on those steps and the input overhead of re-reading a growing conversation on every call. At $0.75 input with a 1M window, high-volume document extraction, agentic search, and classification workloads are exactly where this price cut compounds. Test it before January, because the price doubles after that. DeepSeek V4 Pro goes GA, then raises prices DeepSeek moved V4 Pro to general availability this week with the 0813 checkpoint, ending a preview that began with the April 24 launch. The company updated its API pricing page on August 12 to map the deepseek-v4-pro endpoint to DeepSeek-V4-Pro-0813, and OpenRouter listed the model the same day. There was no blog post and no press release, just a changed model table. Existing integrations keep the same model name and base URL, and the release retains the 1-million-token context window, 384K maximum output, thinking and non-thinking modes, tool calls, and native Responses and Anthropic API compatibility. The benchmark claims are large and unverified. DeepSeek's own table shows broad agent and coding gains over the Pro Preview build, with reported improvements of up to 49.9 percentage points on individual tests, and a Humanity's Last Exam with tools score rising from 48.2 to 60.0. No third-party evaluator has replicated the headline numbers yet. Where independent measurement exists, the picture is more modest: Artificial Analysis scores V4 Pro at 53 on its Intelligence Index, one point above DeepSeek's own near-free V4 Flash at 52 and ten points below Claude Opus 5 at 63. One neutral harness places it second on SWE-bench Verified at 96.40%, behind only Claude Opus 5, while LiveBench ranks it last of seven frontier peers on agentic coding. Strong patch-style coder, weak long-horizon agent. The bigger story is the price reset. Since May, V4 Pro has cost $0.435 per million input tokens on a cache miss, $0.003625 on a cache hit, and $0.87 per million output. This week DeepSeek moved both V4 models to peak and off-peak billing, with V4 Pro at $0.66 input and $1.98 output off-peak and $1.32 and $3.96 at peak. Cache-hit input rises up to 12-fold at peak hours. Even after the increase, per unit of work DeepSeek stays cheap, at roughly $0.06 per completed benchmark task against $2.34 for Claude Opus 5 on the Artificial Analysis measure. But the direction matters: the era of DeepSeek pricing as a loss-leader appears to be ending, and the company itself warns of further increases with no timeline disclosed. Teams that built cost models on DeepSeek's flat rates should rerun the math on their actual traffic hours. GLM-5.3 arrives with CVE receipts Z.ai announced GLM-5.3 on August 14, its new flagship for complex software engineering, long-horizon agentic tasks, and cybersecurity work. Architecturally it follows the same playbook as Grok 4.6: the GLM-5.2 base model, roughly 750 billion parameters, held constant, with all the claimed gains coming from expanded post-training. It supports Low, High, and Max thinking effort and a 1-million-token context window. The distinctive claim is security research capability. Z.ai says the GLM-5 line found 2,436 real vulnerabilities, and unlike most vendor claims, this one has partial external validation: FreeBSD and Red Hat CVE entries credit the model line. That is a new kind of benchmark, one where the scoreboard is public vulnerability databases rather than a lab-controlled harness. The access story is the catch. There are no open weights at launch, a break from Z.ai's history, and no public API for roughly two weeks. Availability starts with GLM Coding Plan subscribers, whose tiers run $18 Lite, $80 Pro, and $168 Max per month, now on a credit system. For a lab that built its reputation on open weights, shipping a closed flagship behind a subscription is a strategic tell worth watching. The rest of the week's releases Four smaller releases filled out the window. Alibaba's Qwen team shipped Qwen3.8-27B on August 14, continuing its fast open-weights cadence. NVIDIA released Nemotron 3.5 Lightning 30B A3B in NVFP4, notable for shipping natively in the 4-bit format its Blackwell hardware accelerates. Dots Studio put out dots3-note Preview, and Mixedbread released Toast 1, a new embedding model. None of these moves the frontier, and all of them widen the menu of small models cheap enough to run everywhere. Tooling: Grok Bot, the Cursor Acquisition, and Public Agent Evals Grok Bot gives agents your logins SpaceXAI opened early beta access to Grok Bot on August 11, one day before Grok 4.6, and the pairing is deliberate. Grok Bo
Comments
No comments yet. Start the discussion.