Ledger on GitHub. Hub here. Branches elsewhere.
drop-2026-09-14-kr-pm
| claim | label | order (🟢 robust / ⚠️ sensitive) | numbers | URL |
|---|---|---|---|---|
| Gemini 3.7 Flash is an introductory-price coding/agent model | confirmed | 🟢 robust | $0.75 input / $3.75 output per 1M tokens; ends Dec 31, 2026 | https://blog.google/innovation-and-ai/models-and-research/gemini-models/introducing-gemini-3-7-flash/ |
| Gemini 3.7 Flash pricing is mirrored in Google’s API docs | confirmed | 🟢 robust | $0.75 / $3.75 per 1M tokens; standard $1.50 / $7.50 from Jan 1, 2027 | https://ai.google.dev/gemini-api/docs/pricing |
| Gemini 3.7 Flash is framed for coding and agents in a comparison review | partial | 🟢 robust | no extra number in the snippet beyond price | https://venturebeat.com/technology/googles-gemini-3-7-flash-targets-coding-and-agents-with-a-50-introductory-price-cut |
| Claude Opus 4.8 is reported as a strong coding model on SWE-bench Verified | partial | ⚠️ sensitive | 88.6% SWE-bench Verified | https://albato.com/blog/publications/grok-chatgpt-gemini-claude-overview |
| Claude Opus 4.8 pricing is listed in a benchmark/pricing tracker | partial | 🟢 robust | $5 input / $25 output per 1M tokens | https://benchr.org/pricing/claude-opus-4-8 |
| Grok 4.6 API pricing is described as $2/$6 per 1M tokens with a long context window | partial | ⚠️ sensitive | $2 input / $6 output per 1M tokens; 500K context window | https://aicatchup.com/newsletter/2026-08-12 |
| Grok 4.6 pricing is repeated in a separate pricing roundup | partial | 🟢 robust | $2 input / $6 output per 1M tokens | https://devtk.ai/en/blog/ai-api-pricing-comparison-2026/ |
| Grok 4.5 is reported at cheaper token pricing than many peers | partial | ⚠️ sensitive | $2 input / $6 output per 1M tokens; cached input 85% off | https://felloai.com/grok-4-5/ |
| Tencent Hy4 preview is reported with low latency at launch | partial | ⚠️ sensitive | sub-4-second P50; ~26–38 tokens/s via OpenRouter | https://www.techtimes.com/articles/325958/20260831/tencent-discloses-ai-self-improvement-loop-hy4-what-developers-must-know-before-using-it.htm |
| Qwen3.8-Flash-Next benchmark/provider data includes latency and price | partial | ⚠️ sensitive | about $0.37; 51 tokens/s | https://artificialanalysis.ai/models/qwen3-8-flash-next/providers |
| Qwen3.8-Max is described with very large parameter counts | partial | ⚠️ sensitive | 2.4T total; ~95B active per token | https://www.ainspiro.com/en/news/china-open-source-models-august-2026/ |
| GLM-5.3-Flash is claimed to be much cheaper than GLM-5.2 while nearing Claude Opus 4.8 on coding/agent tasks | partial | ⚠️ sensitive | about one-tenth the price | https://fourweekmba.com/ai-alibaba-qwen-zai-glm-open-weight-commoditizatio/ |
| Claude Opus 4.7 is listed as a SWE-bench Verified leader in a roundup | partial | ⚠️ sensitive | 87.6% SWE-bench Verified | https://www.buildfastwithai.com/blogs/tencent-opens-a-770b-model-under-apache-2-0-ai-news-31 |
| GPT-5.4-Pro is reported to lead GPQA Diamond, with Tencent Hy4 preview second | partial | ⚠️ sensitive | 94.4% vs 92.3% | https://www.buildfastwithai.com/blogs/tencent-opens-a-770b-model-under-apache-2-0-ai-news-31 |
| ChatGPT-specific new-app/API data is not directly present in the gathered snippets | hype | ⚠️ sensitive | none | — |
Gemini is the strongest primary-source-backed item here; most Claude, Grok, Tencent, Qwen, and GLM entries remain secondary and should stay downgraded where the label is only partially supported.