Ledger on GitHub. Hub here. Branches elsewhere.
drop-2026-09-08-kr-pm
| claim | label | order (🟢 robust / ⚠️ sensitive) | numbers | URL |
|---|---|---|---|---|
| Gemini 3.7 Flash launched with introductory pricing and is positioned for coding/agents | 🟢 confirmed | $0.75/M input; $3.75/M output; 1M context | https://blog.google/innovation-and-ai/models-and-research/gemini-models/introducing-gemini-3-7-flash/ | |
| Gemini 3.7 Flash shows strong coding performance in independent coverage | ⚠️ sensitive | FrontierCode 43.6%; Code Arena 1588 Elo | https://www.datacamp.com/blog/gemini-3-7-flash | |
| Gemini CLI appears as a free coding-agent entry point with published scores | 🟢 confirmed | 70.7% Terminal-Bench 2.1; 80.6% SWE-bench Verified; 54.2% SWE-bench Pro | https://neuralcoretech.com/best-ai-coding-agents-august-2026/ | |
| Grok 4.6 launched with API pricing and long-context tiering | 🟢 confirmed | $2/M input; $0.50/M cached; $6/M output; $4/$1/$12 above 200K | https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-xai-grok-4-6.html | |
| Grok 4.6 coding/agent scores are reported as competitive in launch materials | ⚠️ sensitive | CursorBench 69.9%; DeepSWE 65.9% | https://x.ai/news/grok-4-6 | |
| Grok 4.6 has divergent third-party coding results versus launch claims | ⚠️ sensitive | SWE-bench Verified 95.60%; LiveBench agentic-coding 54.2% | https://codersera.com/blog/grok-4-vs-claude-opus-4-7-vs-gemini-2-5-pro-coding-2026/ | |
| ChatGPT latency is described as sub-500 ms in one comparison, but service reports also show elevated delays/outages | ⚠️ sensitive | under 500 ms; ~0.97–1.05 s first-token checks; ~3-hour outage window | https://www.cnbc.com/2026/08/31/microsoft-outlook-and-openais-chatgpt-work-experience-outages-.html | |
| Claude and ChatGPT entry pricing are tracked in comparison tables, but model-level August deltas are unclear | ⚠️ sensitive | $17–20/mo Pro; $20/mo Plus | https://aitoolsreview.co.uk/insights/claude-vs-chatgpt-gemini-grok | |
| Qwen3.8-Flash is described as an open-weight multimodal MoE release with API pricing | ⚠️ sensitive | 125B params + 51B n-gram; $0.16/M input; $0.47/M output | https://www.requesty.ai/blog/open-weight-frontier-august-2026-glm-qwen-hy4 | |
| GLM-5.3-Flash is described as a low-cost open-weight model with 1M context | ⚠️ sensitive | 320B total; 18B active; 1M context; $0.15/M input; $0.50/M output | https://www.requesty.ai/blog/open-weight-frontier-august-2026-glm-qwen-hy4 | |
| Tencent Hy4 is reported as a very large open-sourced model with hosted access pricing | ⚠️ sensitive | 770B total; 49B active; 1M context; about $0.83/M input | https://www.forbes.com/sites/jonmarkman/2026/08/31/tencent-open-sources-hy4-its-770-billion-parameter-flagship-model/ | |
| New agents this week include Grok Bot with cloud-computer-style always-on workflows | ⚠️ sensitive | beta; $300/mo SuperGrok Heavy | https://www.aipricing.guru/blog/ai-pricing-week-in-review-2026-08-07-to-2026-08-13/ |