Ledger on GitHub. Hub here. Branches elsewhere.
#38: Frontier and Open-Weight Model Landscape Update, August 2026
#38: Frontier and Open-Weight Model Landscape Update, August 2026
Claude Opus 5 launched on July 24, 2026. Pricing for Opus 5 remained unchanged. Performance results were reported to be strong for agentic coding and computer-use tasks.
Claude Opus 5 was subsequently reported to lead agentic and frontier benchmarks in secondary roundups.
Grok 4.6 was listed with frontier-model pricing and a large context window in August 2026 roundups.
Gemini 3.7 Flash was listed as an efficient-tier model with introductory pricing in August 2026 roundups.
GLM-5.3 was reported as a coding and agent flagship. Its API pricing was reported to be unchanged from the prior generation.
GLM-5.3-Flash was reported with an MIT license. It was described as a 320B MoE model with 18B active parameters. It was reported to have a 1M context window and low pricing.
Qwen3.8-Flash-Next was described as an open-weight Qwen4 preview. It included multimodal support and low pricing.
Tencent Hy4 was reported as a 770B open-sourced flagship. It was noted to have 49B active parameters and a 1M context window.
Hy4 hosted access was reported at a very low per-token rate given its model size.
The summer-2026 open-model survey indicated that Chinese labs consistently shipped the largest open models of the year.
A broad AI comparison page summarized model strengths, context limits, and API pricing across major chatbots.
Grok is described as having fast response times in developer-focused comparison coverage.
Claude, ChatGPT, Gemini, and Grok were compared based on plan pricing and feature bundles in a consumer-facing pricing roundup.
GPT-family, Claude, Gemini, and Grok were presented in a broad model-price comparison that included multiple paid tiers.
No gathered source provides strong primary verification for a new Gemini, Claude, ChatGPT, or Grok release this week.
| claim | label | order (๐ข robust / โ ๏ธ sensitive) | numbers | URL |
|---|---|---|---|---|
| Claude Opus 5 launched on July 24, 2026 with unchanged pricing and strong agentic coding/computer-use results. | ๐ข confirmed | ๐ข robust | $5/$25 per 1M tokens; OSWorld 2.0 70.57% | https://www.marktechpost.com/2026/07/24/meet-the-new-claude-opus-5-frontier-class-agentic-coding-and-computer-use-at-unchanged-opus-pricing/ |
| Claude Opus 5 was reported to lead agentic and frontier benchmarks in secondary roundups. | ๐ก partial | โ ๏ธ sensitive | 43.3% Frontier-Bench v0.1; Agentic Index first | https://betterstack.com/community/guides/ai/claude-opus-5/ |
| Grok 4.6 was listed with frontier-model pricing and a large context window in August 2026 roundups. | ๐ก partial | โ ๏ธ sensitive | $2/$6 per 1M tokens; 500K context | https://capitalandcompute.net/blog/new-ai-models-august-2026/ |
| Gemini 3.7 Flash was listed as an efficient-tier model with introductory pricing in August 2026 roundups. | ๐ก partial | โ ๏ธ sensitive | $0.75/$3.75 per 1M tokens | https://capitalandcompute.net/blog/new-ai-models-august-2026/ |
| GLM-5.3 was reported as a coding and agent flagship with unchanged API pricing from the prior generation. | ๐ก partial | โ ๏ธ sensitive | $1.40/$4.40 per 1M tokens | https://themodelgap.com/new-ai-models/2026-08 |
| GLM-5.3-Flash was reported as MIT-licensed, 320B MoE, 18B active, with 1M context and low pricing. | ๐ก partial | โ ๏ธ sensitive | 320B total; 18B active; $0.15/$0.50 per 1M | https://themodelgap.com/new-ai-models/2026-08 |
| Qwen3.8-Flash-Next was described as an open-weight Qwen4 preview with multimodal support and low pricing. | ๐ก partial | โ ๏ธ sensitive | 125B total; 6B active; $0.15/$0.47 per 1M | https://capitalandcompute.net/blog/new-ai-models-august-2026/ |
| Tencent Hy4 was reported as a 770B open-sourced flagship with 49B active parameters and 1M context. | ๐ก partial | โ ๏ธ sensitive | 770B total; 49B active; 1M context | https://www.forbes.com/sites/jonmarkman/2026/08/31/tencent-open-sources-hy4-its-770-billion-parameter-flagship-model/ |
| Hy4 hosted access was reported at a very low per-token rate for a model of its size. | ๐ก partial | โ ๏ธ sensitive | $0.83 per 1M input tokens | https://www.forbes.com/sites/jonmarkman/2026/08/31/tencent-open-sources-hy4-its-770-billion-parameter-flagship-model/ |
| The summer-2026 open-model survey said Chinese labs repeatedly shipped the largest open models of the year. | ๐ก partial | ๐ข robust | 754B to 2.78T monthly ceilings | https://huggingface.co/blog/state-of-open-models-summer-2026 |
| The broad AI comparison page summarized model strengths, context limits, and API pricing across major chatbots. | ๐ก partial | โ ๏ธ sensitive | no single verified total | https://penchan.co/en/ai/compare/ |
| Grok is described as having fast response times in developer-focused comparison coverage. | ๐ก partial | โ ๏ธ sensitive | 1โ4 s; <1 s in one table | https://techlifeadventures.com/post/grok-vs-claude-vs-chatgpt-vs-gemini-developers-2026 |
| Claude, ChatGPT, Gemini, and Grok are compared on plan pricing and feature bundles in a consumer-facing pricing roundup. | ๐ก partial | โ ๏ธ sensitive | $20/$20/$19.99/$30 mo | https://aitoolsreview.co.uk/insights/claude-vs-chatgpt-gemini-grok |
| GPT-family, Claude, Gemini, and Grok are presented in a broad model-price comparison with multiple paid tiers. | ๐ก partial | โ ๏ธ sensitive | $8โ$249.99/mo tiers | https://rg2b.com/chatgpt-vs-claude-vs-gemini-vs-grok-2026/ |
| No gathered source gives strong primary verification for a new Gemini, Claude, ChatGPT, or Grok release this week. |
To do:
- Act: Evaluate Claude Opus 5 for agentic workflows, particularly in coding and computer automation. The reported performance on OSWorld 2.0 suggests this may be a strong candidate for such uses.
- Watch: Monitor the performance and adoption of Tencent Hy4, especially given its reported open-source nature and large parameter count combined with low hosted access pricing. Its large context window also warrants attention.
- Ignore: Claims of new releases from major frontier model providers (Gemini, Claude, ChatGPT, Grok) this week lack strong primary verification. Focus on confirmed launches and model updates.