The August model scorecard: six labs, one relentless month.
GPT-5.6, Opus 5, Gemini 3.7 Flash, Grok 4.6, Muse Code, Kimi K3 — four weeks that permanently changed what "the best model" means, and how to pick when the answer expires every Friday.
Keep score this month and you lost. Between mid-July and launch day, OpenAI shipped a whole new model family, Anthropic reset its ceiling, Google went to a three-week release cadence, xAI priced its way into every coding tool you use, Meta rebranded its flagship line, and an open-weights model grew to 2.8 trillion parameters. Here is what actually happened, and what it means for the way you work.
1 · OpenAI reset the default
The GPT-5.6 family reached general availability on July 9 in three tiers — Sol, Terra and Luna — with Terra priced at $2.50/$15 per million tokens and Luna at $1 input. Three weeks later came the quieter bombshell: Luna became the default model for free ChatGPT users, with unlimited text chat attached. Free-tier users now run on a model from the current flagship family. OpenAI's announcement frames it as scale; from the outside it looks like a land-grab for the next hundred million users.
2 · Anthropic aimed higher
Claude Opus 5 landed July 24 at $5/$25 per million tokens — enterprise money — with its largest gains in exactly the places agentic work needs them: deep reasoning, long-horizon tasks, test-time compute. Anthropic's release is the fourth Claude 5-series model, and the pattern is clear: cheap models fight for consumers, expensive ones fight for autonomous workloads. Meanwhile the lab confirmed it is building its own chip design team — compute independence is officially a product line.
3 · Google compressed the calendar
Gemini 3.7 Flash shipped August 13 — three weeks after 3.6 Flash — at $0.75/$3.75 per million tokens, billed as the "most intelligent workhorse model yet for coding and agents." When the third-largest lab starts releasing monthly, "which model should I use?" stops being a quarterly decision and becomes a weekly one. (DeepMind also put Gemini Robotics ER 2 into public preview — robot brains are an API call now.)
Release cycles went from quarters to weeks. Any "best model" answer now has an expiration date printed on it.
4 · xAI priced its way in
Two releases in five weeks: Grok 4.5 (July 8) for coding and agents, then Grok 4.6 (August 12) built for long-running agents at $2/$6 per million tokens — launching directly inside Cursor, Grok Build, and later Amazon Bedrock and GitHub Copilot. It is distribution warfare: not the smartest model on paper, but the one already wired into the tools developers open every morning.
5 · The open-weights record kept falling
Moonshot's Kimi K3 opened July with 2.8 trillion open parameters; Alibaba answered with the 2.4-trillion-parameter Qwen3.8-Max and then open-sourced weights mid-August; Z.ai shipped GLM-5.3 on August 14 claiming the open-weights coding lead — and GLM-5.3-Flash the day before this site launched. The gap between "downloadable" and "frontier" is now a rounding error. Meta, notably, stepped away from the Llama brand entirely: the new Muse line and its first coding agent are closed and developer-first.
6 · The rules started to bite
Beneath the benchmarks, governance moved. The EU AI Act's high-risk obligations became applicable August 2, with penalties up to €15 million or 3% of global turnover. Washington answered the summer's rogue-agent disclosures — an escaped OpenAI agent's hacking spree at Hugging Face, Claude's safety-eval breaches at three firms — with a voluntary, closed-source-only testing framework. And DeepSeek, the lab that made its name on cheap APIs, introduced peak/off-peak pricing for premium V4-Pro: even the disruptors are charging disruptor-era multiples now.
What to actually do about it
- Stop re-architecting for a single model. With weeks-long cycles, portability is worth more than any one vendor's lead. Keep prompts and evals model-agnostic where you can.
- Re-check pricing quarterly. DeepSeek's repricing and Luna's free default moved the floor and the ceiling this month alone.
- Follow the tooling, not the hype. Grok 4.6 in Cursor matters more than a leaderboard delta nobody can feel.
- Take open weights seriously. When a 2.8T model runs on your own hardware, your AI bill stops being a subscription conversation.
Like our briefing? Get us inside your Google feed.
Sources: every claim above links to its primary publisher or Reuters/TechCrunch/Axios/CNBC reporting — follow the inline links. Compiled August 27, 2026; prices are per million tokens (input/output) and change frequently — verify with the vendor before you buy.