Ledger on GitHub. Hub here. Branches elsewhere.
drop-2026-09-14-kr-close
The strict same-week cross-vendor scoreboard is only partially verifiable from the gathered material; the strongest repeatable signals are pricing and agentic/coding, while search and fresh app/API claims are weaker. Below is a merged, downgraded scoreboard with duplicate-URL avoidance and only claims that appeared in the provided material.
| claim | label | order (🟢 robust / ⚠️ sensitive) | numbers | URL |
|---|---|---|---|---|
| Grok 4.6 kept Grok 4.5’s low pricing and was positioned as a coding/agentic option | confirmed | 🟢 robust | $2 / $6 | https://designforonline.com/the-best-ai-models-so-far-in-2026/ |
| Claude Opus 5 was framed as the premium leader for coding and long-horizon agentic work | confirmed | 🟢 robust | $5 / $25 | https://www.stackspend.app/blog/model-selection/best-coding-models-august-2026 |
| Gemini 3.7 Flash launched with lower pricing than the prior rate and was presented as a cheaper flagship tier | confirmed | 🟢 robust | $0.75 / $3.75 per 1M tokens | https://singhajit.com/dev-weekly/2026/aug-10-16/gemini-3-7-flash-grok-4-6-patch-tuesday-anthropic-2t-ipo-chaindrop/ |
| GPT-5.6 Sol appeared in a comparison where its agentic performance trailed Claude Opus 5 and Grok 4.6 on some rounds | partial | ⚠️ sensitive | 73% resolved; $8.39 measured cost/task | https://rohitai.com/blog/best-ai-models-2026-openai-anthropic-google-xai-deepseek |
| Gemini 3.7 Flash showed strong throughput and low first-answer latency in a scorecard | partial | ⚠️ sensitive | 340.07 tok/s; 9.83s first answer; $2.18 measured cost/task | https://rohitai.com/blog/best-ai-models-2026-openai-anthropic-google-xai-deepseek |
| Grok 4.6 was described as beating GPT-5.6 Sol on coding-related metrics in a roundup | partial | ⚠️ sensitive | Intelligence 60.92; Agentic 58.68; Terminal-Bench 88.39% | https://rohitai.com/blog/best-ai-models-2026-openai-anthropic-google-xai-deepseek |
| Claude Opus 5 was reported as the top model on an intelligence-style scoreboard in secondary coverage | partial | ⚠️ sensitive | Intelligence Index 63; Agentic Index 55.3 | https://felloai.com/best-ai-models/ |
| Grok 4.5 was listed with low-latency/high-throughput positioning on a comparison page | partial | ⚠️ sensitive | 80 TPS; 500k context; $2 / $6 | https://developersdigest.tech/topics/ai-models |
| Gemini 3.1 Pro was listed in pricing trackers as the cheaper flagship among big-brand alternatives | partial | ⚠️ sensitive | $2.00 / $12.00 per 1M tokens | https://compareai.today/blog/ai-api-pricing-august-2026 |
| AI search APIs for agents were being compared as a new operational lane, with low-latency search and MCP-style access highlighted | partial | ⚠️ sensitive | ~0.5s to ~3s; per-request pricing varied | https://techsy.io/en/blog/best-ai-search-apis-2026 |
| A search/API provider page highlighted agentic search endpoints and cited latency-oriented tiers for production use | partial | ⚠️ sensitive | /search/fast, /search/content, /answer, /agentic; 1s to 30s+ | https://keirolabs.cloud/blogs/guide/best-ai-search-apis-for-agents-aug-2026 |
The table stays on the scoreboard lane: model pricing, coding, agents, latency, and search/API signals. Items marked sensitive are retained but downgraded because the source chain is mostly secondary or tracker-style rather than primary docs.