dropkit.contents

Ledger on GitHub. Hub here. Branches elsewhere.

← Drops · 2026-09-05

drop-2026-09-05-us-lunch

August 2026 AI Model Release Summary

Claude Opus 5 led the August 2026 Artificial Analysis Intelligence Index. This model carried a premium price. Its API was priced at $5 in per 1M tokens and $25 out per 1M tokens. The AA Index score for Claude Opus 5 was 63. Claude Opus 5 was also positioned as the coding and knowledge-work leader among listed assistants. Its API price remained high. The AA Intelligence Index score for Claude Opus 5 in this context was 60.7.

GPT-5.6 Sol sat below Claude Opus 5 on the same August 2026 scoreboard. This model had a higher output price than Claude Opus 5. The AA Index score for GPT-5.6 Sol was 61. Its API was priced at $5 in per 1M tokens and $30 out per 1M tokens.

Grok 4.6 launched with competitive price tiers. It demonstrated low latency on its model page. Task pricing for Grok 4.6 was $0.25, $0.78, $0.94, and $1.23. The time to first token was 8.25 seconds. Grok 4.6 shipped in August 2026. It was marketed for agentic and long-running tasks. Its token price was lower than Claude Opus.

DeepSeek V4 Pro 0813 changed to peak and off-peak API pricing. Peak pricing was $1.32 in per 1M tokens and $3.96 out per 1M tokens. Off-peak pricing was $0.66 in per 1M tokens and $1.98 out per 1M tokens. DeepSeek V4 Pro was also reported as a 1.6T-parameter MoE. It achieved 96.40% on SWE-bench Verified.

GLM-5.3 showed a large jump on agent and coding benchmarks after launch-side post-training. Terminal-Bench 3.0 scores increased from 4.6 to 28.3. DeepSWE v1.1 scores increased from 46.2 to 66.9. CyberGym scores were 84.5%. GLM-5.3 was described as an open-weights coding and cyber model. It showed a large Terminal-Bench improvement, with a score of 28.3. Its CyberGym score was 84.5%.

Tencent Hy4 was described as a 770B flagship model. It featured 49B active parameters and a 1M-token context window. The estimated cost was approximately $0.83 per 1M input tokens.

Qwen 3.8-Flash was presented as a multimodal MoE open-weight release. It offered very low API pricing. The model had 125B parameters and a 51B n-gram component. Its API pricing was $0.16 in per 1M tokens and $0.47 out per 1M tokens.

Qwen 3.8-Max was positioned near the frontier. It was described as being below the very top closed models. Its AA Index score was 58. The API pricing was $2 in per 1M tokens and $6 out per 1M tokens.

BenchLM’s C-Eval leaderboard put DeepSeek V4 Pro Base slightly ahead of Qwen3.5 397B. The scores were 93.1% for DeepSeek V4 Pro Base and 93.0% for Qwen3.5 397B. BenchLM’s best Chinese models page had Kimi K3 leading. It had a narrow confidence band, with a score of 78.3 and a 90% interval of 76.05–80.58.

claimlabelorder (🟢 robust / ⚠️ sensitive)numbersURL
Claude Opus 5 led the August 2026 Artificial Analysis Intelligence Index and carried a premium price.🟢 confirmed🟢 robustAA Index 63; $5 in / $25 out per 1M tokenshttps://www.anthropic.com/news/claude-opus-5
Claude Opus 5 was positioned as the coding/knowledge-work leader among the listed assistants, but its API price stayed high.🟢 confirmed🟢 robust$5/$25 per 1M tokens; AA Intelligence Index 60.7https://www.anthropic.com/news/claude-opus-5
Grok 4.6 was launched with competitive price tiers and low latency on the model page.🟡 partial⚠️ sensitive$0.25 / $0.78 / $0.94 / $1.23 task pricing; 8.25s time to first tokenhttps://artificialanalysis.ai/models/releases/grok-4-6
Grok 4.6 shipped in August 2026 and was marketed for agentic/long-running tasks, with a lower token price than Claude Opus🟢 confirmed🟢 robustnot in sourcesnot in sources
DeepSeek V4 Pro 0813 changed to peak/off-peak API pricing.🟢 confirmed🟢 robust$1.32 in / $3.96 out peak; $0.66 in / $1.98 out off-peak per 1M tokenshttps://www.reuters.com/world/china/deepseek-launches-v4-pro-at-prices-up-to-14-times-higher-than-openais-least-expensive-model-2026-08-14/

Act: Monitor the sustained performance and pricing of Claude Opus 5, especially its position in coding and knowledge work. Watch: Observe the competitive positioning of Grok 4.6, particularly its performance on agentic and long-running tasks given its token price compared to Claude Opus. Ignore: Specific benchmark scores that lack corroborating details on methodology or wider industry acceptance, as these can be sensitive and subject to change.