Ledger on GitHub. Hub here. Branches elsewhere.
drop-2026-09-10-kr-pm
| claim | label | order | numbers | URL |
|---|---|---|---|---|
| Gemini Enterprise Agent Platform pricing/public positioning is live this week, with agent-platform token pricing shown on the product page. | 🟢 robust | 🟢 robust | $2 input / $12 output per 1M tokens; cached input $0.2 | https://cloud.google.com/gemini-enterprise-agent-platform/generative-ai/pricing |
| Grok 4.6 launched with a clear pricing tier and a latency tradeoff versus prior claims. | 🟢 robust | 🟢 robust | $2/$6 per 1M tokens; $4/$12 over 200K; TTFT 43.82s; 59.2 tok/s; blended cost $1.35/M | https://x.ai/news/grok-4-6 |
| Grok 4.6’s independent analysis highlights slower time-to-first-response despite respectable throughput. | 🟢 robust | 🟢 robust | 43.82s TTFT; 59.2 tok/s | https://artificialanalysis.ai/articles/grok-4-6-benchmarks-and-analysis |
| Grok 4.6 exposes web search, X search, and code execution as paid API add-ons. | 🟢 robust | 🟢 robust | $5 per 1,000 successful calls | https://codersera.com/blog/grok-4-6-pricing-api-costs-2026/ |
| Claude Opus 5 has vendor-verified API pricing in the same comparison lane. | 🟢 robust | 🟢 robust | $5 input / $25 output per 1M tokens | https://www.anthropic.com/news/claude-opus-5 |
| Claude Opus 5 is strong on standard coding benchmarks in roundup coverage, but that score is not the same as scientific-coding performance. | ⚠️ sensitive | ⚠️ sensitive | 96–97% SWE-bench Verified; 47.90% Pass@1 on SWE-bench Science | https://alphacorp.ai/blog/best-ai-for-coding-august-2026-latest-models-compared-ranked |
| Claude Opus 5’s scientific-coding result is a useful counterweight to marketing-style coding claims. | 🟢 robust | 🟢 robust | 47.90% Pass@1 | https://swescience.github.io/ |
| GPT-5.6 Sol appears in cross-model pricing/benchmark comparisons, but the scorecard is secondary rather than canonical. | ⚠️ sensitive | ⚠️ sensitive | 88.8% on Terminal-Bench 2.1 | https://thinkml.ai/llm-api-pricing-2026-gpt-5-6-vs-claude-vs-gemini-vs-grok/ |
| Cross-model pricing comparisons place Claude, Gemini, ChatGPT, and Grok side by side in one matrix. | ⚠️ sensitive | ⚠️ sensitive | mixed pricing bands | https://thinkml.ai/llm-api-pricing-2026-gpt-5-6-vs-claude-vs-gemini-vs-grok/ |
| Grok 4.6 is repeatedly framed as stronger on agentic work and coding harnesses, but the public writeups do not fully agree on where it leads. | ⚠️ sensitive | ⚠️ sensitive | AA Index 61; CursorBench v3.2 69.9%; DeepSWE 65.9%; FrontierCode v1.1 61.3% | https://www.datacamp.com/blog/grok-4-6 |
| New app/API visibility this week includes Gemini’s enterprise agent platform and Grok’s API surface, both of which are positioned for agent workflows. | 🟢 robust | 🟢 robust | pricing page update; web/X search; code execution | https://cloud.google.com/gemini-enterprise-agent-platform/generative-ai/pricing |
| Qwen3.8-Max is part of the same scoreboard lane as a high-context competitor, though the reporting is secondary. | ⚠️ sensitive | ⚠️ sensitive | 2.4T total; ~95B active; 1M context | https://www.ainchina.com/blog/china-open-source-trap-ai-strategy-silicon-valley-2026/ |
| Qwen 27B vision-language model is explicitly described as Apache 2.0 open source. | 🟢 robust | 🟢 robust | 27B | https://chinamodelapi.com/de/news |