Ledger on GitHub. Hub here. Branches elsewhere.
A Review of Eleven Models and Their Positions
A Review of Eleven Models and Their Positions
Gemini 3.7 Flash was released with API pricing and a focus on coding and agent workflows. The pricing was listed at $0.75 input and $3.75 output per 1M tokens. This model included a 1,048,576-token context window.
Claude Opus 5 maintained a high position on coding-agent benchmarks. It achieved 89.1% on Terminal-Bench 2.1.
GPT-5.6 Sol surpassed Claude Opus 5 on the same benchmark. It registered 89.5% on Terminal-Bench 2.1. GPT-5.6 Sol was also used to establish pricing comparisons for ChatGPT/Codex, with ChatGPT Plus via Codex noted at $20/month.
Grok 4.6 was positioned for long-running agents and interactive or visual tasks. No verified benchmark was available in the source for this claim.
ChatGPT GPT-5.6 Sol experienced reported price reductions in late August. The new pricing was $4 per 1M tokens for input and $20 per 1M tokens for output.
DeepSeek V4 Pro launched with competitive API pricing and claims related to agent capabilities. Its pricing was $0.435 input and $0.87 output per 1M tokens. It showed a 2.8% average gap on 9 tasks.
Tencent Hy4 was reported as an open-sourced flagship model with 770B parameters.
Qwen3.8-Flash was described as an open-weight multimodal MoE preview. Its specifications included 125B total parameters and 51B N-gram. API prices were stated as $0.16 input and $0.47 output per 1M tokens.
Chinese open-model activity was a significant factor in open-weight releases during summer 2026. This period saw 178 releases with over 20B parameters, with 59% licensed under Apache 2.0 and 22% under MIT.
Chinese monthly open-model parameter ceilings reportedly surpassed those in the U.S. for most of 2026. The monthly ceilings ranged from 754B to 2.78T parameters.
| claim | label | order (🟢 robust / ⚠️ sensitive) | numbers | URL |
|---|---|---|---|---|
| Gemini 3.7 Flash launched with introductory API pricing and stronger coding/agent positioning | 🟢 confirmed | 🟢 robust | $0.75 input / $3.75 output per 1M | https://blog.google/innovation-and-ai/models-and-research/gemini-models/introducing-gemini-3-7-flash/ |
| Gemini 3.7 Flash was framed as built for complex coding and agentic workflows | 🟢 confirmed | 🟢 robust | 1,048,576-token context window | https://blog.google/innovation-and-ai/models-and-research/gemini-models/introducing-gemini-3-7-flash/ |
| Claude Opus 5 stayed near the top on coding-agent benchmarks | 🟢 confirmed | 🟢 robust | 89.1% Terminal-Bench 2.1 | https://www.neuralcoretech.com/best-ai-coding-agents-august-2026/ |
| GPT-5.6 Sol edged Claude Opus 5 on Terminal-Bench 2.1 | 🟢 confirmed | 🟢 robust | 89.5% Terminal-Bench 2.1 | https://www.neuralcoretech.com/best-ai-coding-agents-august-2026/ |
| GPT-5.6 Sol was also used as a pricing comparison anchor for ChatGPT/Codex | 🟢 confirmed | 🟢 robust | $20/mo ChatGPT Plus via Codex | https://www.neuralcoretech.com/best-ai-coding-agents-august-2026/ |
| Grok 4.6 was positioned around long-running agents and interactive/visual work | ⚠️ partial | ⚠️ sensitive | no verified benchmark in source | https://nextgenrationaiworld1.blogspot.com/2026/08/whats-new-in-chatgpt-gemini-claude-and.html?m=1 |
| ChatGPT GPT-5.6 Sol was reported as cheaper after late-August price cuts | ⚠️ partial | ⚠️ sensitive | $4 / $20 per 1M tokens | https://kafkai.ai/articles/ai-technology/ai-models-july-august-2026-overview/ |
| DeepSeek V4 Pro launched with aggressive API pricing and agent claims | ⚠️ partial | ⚠️ sensitive | $0.435 input / $0.87 output per 1M; 2.8% average gap on 9 tasks | https://hyper.ai/en/stories/f5fe72f90680d52cb765bbb8dd09e77d |
| Tencent Hy4 was reported as a 770B open-sourced flagship model | ⚠️ partial | ⚠️ sensitive | 770B parameters | https://www.forbes.com/sites/jonmarkman/2026/08/31/tencent-open-sources-hy4-its-770-billion-parameter-flagship-model/ |
| Qwen3.8-Flash was described as an open-weight multimodal MoE preview with low API prices | ⚠️ partial | ⚠️ sensitive | 125B total; 51B N-gram; $0.16 / $0.47 per M tokens | https://www.requesty.ai/blog/open-weight-frontier-august-2026-glm-qwen-hy4 |
| Chinese open-model activity dominated summer 2026 open-weight releases | 🟢 confirmed | 🟢 robust | 178 releases above 20B; 59% Apache 2.0; 22% MIT | https://huggingface.co/blog/state-of-open-models-summer-2026 |
| Chinese monthly open-model ceilings reportedly exceeded U.S. ceilings through most of 2026 | ⚠️ partial | ⚠️ sensitive | 754B to 2.78T parameter monthly ceilings | https://huggingface.co/blog/state-of-open-models-summer-2026 |
Act: Evaluate Gemini 3.7 Flash for agentic workflows, considering its context window and API pricing. Watch: Monitor the continued performance of GPT-5.6 Sol and Claude Opus 5 on coding benchmarks for shifts in leadership. Ignore: Claims regarding Grok 4.6’s positioning lacking verified benchmark data.