dropkit.contents

Ledger on GitHub. Hub here. Branches elsewhere.

← Drops · 2026-09-07

A Review of Eleven Models and Their Positions

A Review of Eleven Models and Their Positions

Gemini 3.7 Flash was released with API pricing and a focus on coding and agent workflows. The pricing was listed at $0.75 input and $3.75 output per 1M tokens. This model included a 1,048,576-token context window.

Claude Opus 5 maintained a high position on coding-agent benchmarks. It achieved 89.1% on Terminal-Bench 2.1.

GPT-5.6 Sol surpassed Claude Opus 5 on the same benchmark. It registered 89.5% on Terminal-Bench 2.1. GPT-5.6 Sol was also used to establish pricing comparisons for ChatGPT/Codex, with ChatGPT Plus via Codex noted at $20/month.

Grok 4.6 was positioned for long-running agents and interactive or visual tasks. No verified benchmark was available in the source for this claim.

ChatGPT GPT-5.6 Sol experienced reported price reductions in late August. The new pricing was $4 per 1M tokens for input and $20 per 1M tokens for output.

DeepSeek V4 Pro launched with competitive API pricing and claims related to agent capabilities. Its pricing was $0.435 input and $0.87 output per 1M tokens. It showed a 2.8% average gap on 9 tasks.

Tencent Hy4 was reported as an open-sourced flagship model with 770B parameters.

Qwen3.8-Flash was described as an open-weight multimodal MoE preview. Its specifications included 125B total parameters and 51B N-gram. API prices were stated as $0.16 input and $0.47 output per 1M tokens.

Chinese open-model activity was a significant factor in open-weight releases during summer 2026. This period saw 178 releases with over 20B parameters, with 59% licensed under Apache 2.0 and 22% under MIT.

Chinese monthly open-model parameter ceilings reportedly surpassed those in the U.S. for most of 2026. The monthly ceilings ranged from 754B to 2.78T parameters.

claimlabelorder (🟢 robust / ⚠️ sensitive)numbersURL
Gemini 3.7 Flash launched with introductory API pricing and stronger coding/agent positioning🟢 confirmed🟢 robust$0.75 input / $3.75 output per 1Mhttps://blog.google/innovation-and-ai/models-and-research/gemini-models/introducing-gemini-3-7-flash/
Gemini 3.7 Flash was framed as built for complex coding and agentic workflows🟢 confirmed🟢 robust1,048,576-token context windowhttps://blog.google/innovation-and-ai/models-and-research/gemini-models/introducing-gemini-3-7-flash/
Claude Opus 5 stayed near the top on coding-agent benchmarks🟢 confirmed🟢 robust89.1% Terminal-Bench 2.1https://www.neuralcoretech.com/best-ai-coding-agents-august-2026/
GPT-5.6 Sol edged Claude Opus 5 on Terminal-Bench 2.1🟢 confirmed🟢 robust89.5% Terminal-Bench 2.1https://www.neuralcoretech.com/best-ai-coding-agents-august-2026/
GPT-5.6 Sol was also used as a pricing comparison anchor for ChatGPT/Codex🟢 confirmed🟢 robust$20/mo ChatGPT Plus via Codexhttps://www.neuralcoretech.com/best-ai-coding-agents-august-2026/
Grok 4.6 was positioned around long-running agents and interactive/visual work⚠️ partial⚠️ sensitiveno verified benchmark in sourcehttps://nextgenrationaiworld1.blogspot.com/2026/08/whats-new-in-chatgpt-gemini-claude-and.html?m=1
ChatGPT GPT-5.6 Sol was reported as cheaper after late-August price cuts⚠️ partial⚠️ sensitive$4 / $20 per 1M tokenshttps://kafkai.ai/articles/ai-technology/ai-models-july-august-2026-overview/
DeepSeek V4 Pro launched with aggressive API pricing and agent claims⚠️ partial⚠️ sensitive$0.435 input / $0.87 output per 1M; 2.8% average gap on 9 taskshttps://hyper.ai/en/stories/f5fe72f90680d52cb765bbb8dd09e77d
Tencent Hy4 was reported as a 770B open-sourced flagship model⚠️ partial⚠️ sensitive770B parametershttps://www.forbes.com/sites/jonmarkman/2026/08/31/tencent-open-sources-hy4-its-770-billion-parameter-flagship-model/
Qwen3.8-Flash was described as an open-weight multimodal MoE preview with low API prices⚠️ partial⚠️ sensitive125B total; 51B N-gram; $0.16 / $0.47 per M tokenshttps://www.requesty.ai/blog/open-weight-frontier-august-2026-glm-qwen-hy4
Chinese open-model activity dominated summer 2026 open-weight releases🟢 confirmed🟢 robust178 releases above 20B; 59% Apache 2.0; 22% MIThttps://huggingface.co/blog/state-of-open-models-summer-2026
Chinese monthly open-model ceilings reportedly exceeded U.S. ceilings through most of 2026⚠️ partial⚠️ sensitive754B to 2.78T parameter monthly ceilingshttps://huggingface.co/blog/state-of-open-models-summer-2026

Act: Evaluate Gemini 3.7 Flash for agentic workflows, considering its context window and API pricing. Watch: Monitor the continued performance of GPT-5.6 Sol and Claude Opus 5 on coding benchmarks for shifts in leadership. Ignore: Claims regarding Grok 4.6’s positioning lacking verified benchmark data.