Ledger on GitHub. Hub here. Branches elsewhere.
drop-2026-09-05-us-lunch
August 2026 AI Model Release Summary
Claude Opus 5 led the August 2026 Artificial Analysis Intelligence Index. This model carried a premium price. Its API was priced at $5 in per 1M tokens and $25 out per 1M tokens. The AA Index score for Claude Opus 5 was 63. Claude Opus 5 was also positioned as the coding and knowledge-work leader among listed assistants. Its API price remained high. The AA Intelligence Index score for Claude Opus 5 in this context was 60.7.
GPT-5.6 Sol sat below Claude Opus 5 on the same August 2026 scoreboard. This model had a higher output price than Claude Opus 5. The AA Index score for GPT-5.6 Sol was 61. Its API was priced at $5 in per 1M tokens and $30 out per 1M tokens.
Grok 4.6 launched with competitive price tiers. It demonstrated low latency on its model page. Task pricing for Grok 4.6 was $0.25, $0.78, $0.94, and $1.23. The time to first token was 8.25 seconds. Grok 4.6 shipped in August 2026. It was marketed for agentic and long-running tasks. Its token price was lower than Claude Opus.
DeepSeek V4 Pro 0813 changed to peak and off-peak API pricing. Peak pricing was $1.32 in per 1M tokens and $3.96 out per 1M tokens. Off-peak pricing was $0.66 in per 1M tokens and $1.98 out per 1M tokens. DeepSeek V4 Pro was also reported as a 1.6T-parameter MoE. It achieved 96.40% on SWE-bench Verified.
GLM-5.3 showed a large jump on agent and coding benchmarks after launch-side post-training. Terminal-Bench 3.0 scores increased from 4.6 to 28.3. DeepSWE v1.1 scores increased from 46.2 to 66.9. CyberGym scores were 84.5%. GLM-5.3 was described as an open-weights coding and cyber model. It showed a large Terminal-Bench improvement, with a score of 28.3. Its CyberGym score was 84.5%.
Tencent Hy4 was described as a 770B flagship model. It featured 49B active parameters and a 1M-token context window. The estimated cost was approximately $0.83 per 1M input tokens.
Qwen 3.8-Flash was presented as a multimodal MoE open-weight release. It offered very low API pricing. The model had 125B parameters and a 51B n-gram component. Its API pricing was $0.16 in per 1M tokens and $0.47 out per 1M tokens.
Qwen 3.8-Max was positioned near the frontier. It was described as being below the very top closed models. Its AA Index score was 58. The API pricing was $2 in per 1M tokens and $6 out per 1M tokens.
BenchLM’s C-Eval leaderboard put DeepSeek V4 Pro Base slightly ahead of Qwen3.5 397B. The scores were 93.1% for DeepSeek V4 Pro Base and 93.0% for Qwen3.5 397B. BenchLM’s best Chinese models page had Kimi K3 leading. It had a narrow confidence band, with a score of 78.3 and a 90% interval of 76.05–80.58.
| claim | label | order (🟢 robust / ⚠️ sensitive) | numbers | URL |
|---|---|---|---|---|
| Claude Opus 5 led the August 2026 Artificial Analysis Intelligence Index and carried a premium price. | 🟢 confirmed | 🟢 robust | AA Index 63; $5 in / $25 out per 1M tokens | https://www.anthropic.com/news/claude-opus-5 |
| Claude Opus 5 was positioned as the coding/knowledge-work leader among the listed assistants, but its API price stayed high. | 🟢 confirmed | 🟢 robust | $5/$25 per 1M tokens; AA Intelligence Index 60.7 | https://www.anthropic.com/news/claude-opus-5 |
| Grok 4.6 was launched with competitive price tiers and low latency on the model page. | 🟡 partial | ⚠️ sensitive | $0.25 / $0.78 / $0.94 / $1.23 task pricing; 8.25s time to first token | https://artificialanalysis.ai/models/releases/grok-4-6 |
| Grok 4.6 shipped in August 2026 and was marketed for agentic/long-running tasks, with a lower token price than Claude Opus | 🟢 confirmed | 🟢 robust | not in sources | not in sources |
| DeepSeek V4 Pro 0813 changed to peak/off-peak API pricing. | 🟢 confirmed | 🟢 robust | $1.32 in / $3.96 out peak; $0.66 in / $1.98 out off-peak per 1M tokens | https://www.reuters.com/world/china/deepseek-launches-v4-pro-at-prices-up-to-14-times-higher-than-openais-least-expensive-model-2026-08-14/ |
Act: Monitor the sustained performance and pricing of Claude Opus 5, especially its position in coding and knowledge work. Watch: Observe the competitive positioning of Grok 4.6, particularly its performance on agentic and long-running tasks given its token price compared to Claude Opus. Ignore: Specific benchmark scores that lack corroborating details on methodology or wider industry acceptance, as these can be sensitive and subject to change.