Ledger on GitHub. Hub here. Branches elsewhere.
drop-2026-09-12-kr-close
| claim | label | order (🟢 robust / ⚠️ sensitive) | numbers | URL |
|---|---|---|---|---|
| Gemini 3.7 Flash is positioned as a low-cost coding/agent workhorse. | 🟢 confirmed | 🟢 robust | $0.75 input / $3.75 output per 1M | https://blog.google/innovation-and-ai/models-and-research/gemini-models/introducing-gemini-3-7-flash/ |
| Claude Sonnet 5 is priced in the mid-tier API band. | 🟢 confirmed | 🟢 robust | $2 input / $10 output per 1M | https://www.anthropic.com/news/claude-sonnet-5 |
| Grok 4.6 is priced aggressively for API usage. | ⚠️ partial | ⚠️ sensitive | $2 input / $6 output per 1M | https://codersera.com/blog/grok-4-6-pricing-api-costs-2026/ |
| Grok 4.6 includes coding-agent positioning and standard context scale. | ⚠️ partial | ⚠️ sensitive | 256K context | https://codersera.com/blog/grok-4-6-pricing-api-costs-2026/ |
| GPT-5.6 Sol is referenced as a higher-priced benchmark competitor. | ⚠️ partial | ⚠️ sensitive | $5 input / $30 output per 1M | https://aitoolsrecap.com/Blog/ai-price-war-august-2026-openai-anthropic-deepseek |
| Gemini 3.1 Flash leads the cited latency table. | 🟢 confirmed | 🟢 robust | 180 ms TTFT; 210 tok/s | https://railwail.com/al/reports/state-of-ai-apis-2026 |
| Claude 4.7 Sonnet is slower than Gemini 3.1 Flash in the cited latency table. | 🟢 confirmed | 🟢 robust | 320 ms TTFT; 84 tok/s | https://railwail.com/al/reports/state-of-ai-apis-2026 |
| xAI Grok 4.3 lags Gemini 3.1 Flash and Claude 4.7 Sonnet on the cited latency table. | 🟢 confirmed | 🟢 robust | 410 ms TTFT; 68 tok/s | https://railwail.com/al/reports/state-of-ai-apis-2026 |
| Tencent opened a 770B model under Apache 2.0 with benchmark scores. | ⚠️ partial | ⚠️ sensitive | 770B; 92.3 GPQA Diamond; 65.7 SWE-bench Pro | https://www.buildfastwithai.com/blogs/tencent-opens-a-770b-model-under-apache-2-0-ai-news-aug-31 |
| Qwen3.8-Max launched with agentic/coding benchmark claims and API pricing. | ⚠️ partial | ⚠️ sensitive | 86.6 Terminal-Bench 2.1; 67.7 SWE-bench Pro; $2 / $6 per 1M | https://www.techtimes.com/articles/322773/20260803/qwen38-max-debuts-arenaai-qwenwork-brings-china-state-law-risk-enterprise-workflows.htm |
| GLM-5.3-Flash was described as a low-cost open-weight model with very large context. | ⚠️ partial | ⚠️ sensitive | 320B MoE; 18B active; 1M context; $0.15 / $0.50 per 1M | https://www.requesty.ai/blog/open-weight-frontier-august-2026-glm-qwen-hy4 |
| Tencent Hy4 preview was reported to beat GPT-5.6 Sol on some tests. | ⚠️ partial | ⚠️ sensitive | 1-bit quantized variant around 213GB | https://gigazine.net/gsc_news/en/20260831-hy4-preview-tencent/ |
| Gemini 3.7 Flash was positioned as a cheap workhorse for coding and agent workloads. | 🟢 confirmed | 🟢 robust | $0.75 input / 1M | https://blog.google/innovation-and-ai/models-and-research/gemini-models/introducing-gemini-3-7-flash/ |
| Claude Pro was priced at $20/month, with separate monthly credits for higher tiers. | ⚠️ partial | ⚠️ sensitive | $20/mo; $100/mo; $200/mo | https://tech-insider.org/chatgpt-vs-gemini-vs-claude-pro-2026/ |
| Grok 4.5 was highlighted more for price than benchmark dominance. | ⚠️ partial | ⚠️ sensitive | $2 input / $6 output per 1M | https://tech-insider.org/chatgpt-vs-gemini-vs-claude-pro-2026/ |
| Claude Opus 5 / Sonnet 5 / Haiku 4.5 had published base pricing tiers. | � |