On July 31, OpenAI cut GPT-5.6 Luna 80% while DeepSeek shipped V4-Flash-0731. Analysts now rank by intelligence per dollar. Cost per task — not token price — is what teams should model before adopting any agent.
Claude Code hit $2.5B annualized revenue while Codex sits near $1B. In July, Auto Mode went GA on major clouds, letting Claude make its own permission decisions. The win is economics; the risk is permissions.
Google's Gemini 3.6 Flash (GA, July 21) cuts output tokens 17% and output price 16.7%, brings Computer Use to the production API. Token efficiency is now a first-class cost metric for agent workloads.
Zhipu's GLM-5.2 (open-source June 17, MIT, 744B/40B, 1M context) ranks fifth globally in weekly token usage. Edge: long-horizon tasks and Day-0 domestic-chip adaptation — built for Chinese silicon, not just benchmarks.
Alibaba's Qwen3.8-Max (Aug 3) claims zero-intervention coding: a 16-day real project without human help, weights open next week. Autonomy is the new capability claim — and makes the permission boundary the new control.
Meta launched Muse Code (beta) on August 5 — a terminal coding agent on the Muse Spark 1.2 model, priced under 40% of Claude Code and Codex. What engineering teams should verify before adopting.
Moonshot AI released Kimi K3 as open weights in July 2026. What the model card claims, what its custom license requires, and how self-hosting teams should evaluate it.
Developers feel faster with AI, yet DORA 2024 links higher adoption to small drops in delivery throughput and stability—here is how to roll out agents safely.