Ledger on GitHub. Hub here. Branches elsewhere.
drop-2026-09-11-kr-close
Scoreboard merge only: Gemini / Claude / ChatGPT / Grok plus this-week apps and APIs. Dual-pass same label is ๐ข robust; one-pass or flipped labels are โ ๏ธ sensitive.
| claim | label | order | numbers | URL |
|---|---|---|---|---|
| Grok 4.6 has the lowest time to first answer token in the cited leaderboard snapshot | ๐ข confirmed | โ ๏ธ sensitive | 7.02s | https://artificialanalysis.ai/models/releases/grok-4-6 |
| Grok 4.6 has the lowest cost per task in the cited leaderboard snapshot | ๐ข confirmed | โ ๏ธ sensitive | $0.48 | https://artificialanalysis.ai/models/releases/grok-4-6 |
| GPT-5.4 mini is the fastest OpenAI model in the cited tracker | ๐ข confirmed | โ ๏ธ sensitive | 0.72s | https://artificialanalysis.ai/providers/openai |
| GPT-5 nano is the cheapest OpenAI model in the cited tracker | ๐ข confirmed | โ ๏ธ sensitive | $0.05 / 1M tokens | https://artificialanalysis.ai/providers/openai |
| Gemini 2.5 Flash-Lite is the lowest-latency model in the cited cross-provider leaderboard | ๐ข confirmed | โ ๏ธ sensitive | 0.31s | https://artificialanalysis.ai/models |
| Gemini 3.7 Flash shows very low latency in third-party model tracking | ๐ข confirmed | โ ๏ธ sensitive | 0.66s TTFA | https://artificialanalysis.ai/models/releases/gemini-3-7-flash |
| Gemini 3.7 Flash pricing is reported at low API rates | ๐ข confirmed | โ ๏ธ sensitive | $0.75 / $3.75 per 1M tokens | https://www.eesel.ai/blog/gemini-3-7-flash-review |
| Gemini 3.7 Flash latency is also reported as highly variable in community discussion | ๐ก partial | โ ๏ธ sensitive | 2.1โ26.6s | https://discuss.ai.google.dev/t/high-latency-for-ai-studio-gemini-3-7-flash/180447 |
| Claude Sonnet 5 is listed at $2 / $10 in August 2026 pricing tables | ๐ก partial | ๐ข robust | $2 / $10 | https://compareai.today/blog/ai-api-pricing-august-2026 |
| Gemini 3.1 Pro Preview has listed pricing in the cited August 2026 cost table | ๐ก partial | โ ๏ธ sensitive | $2 / $12 | https://modelrefs.com/news/state-of-the-api-2026-08/ |
| xAI Grok 4.6 API pricing is stated in the cited August 2026 pricing article | ๐ก partial | โ ๏ธ sensitive | $2 / $0.50 / $6 | https://codersera.com/blog/grok-4-6-pricing-api-costs-2026/ |
| ChatGPT/GPT-5.5 is cited as premium in pricing comparisons | ๐ก partial | โ ๏ธ sensitive | $5 / $30 per 1M tokens | https://tech-insider.org/grok-vs-chatgpt-vs-gemini-2026/ |
| ChatGPT/GPT-5.6 is mentioned in cross-model score comparisons versus GLM-5.3 | ๐ก partial | โ ๏ธ sensitive | score context only | https://gigazine.net/gsc_news/en/20260829-glm-5-3-open/ |
| Claude 3 Haiku appears in latency snapshots | ๐ก partial | โ ๏ธ sensitive | 553ms avg total; 514ms avg TTFT | https://www.ailatency.com/reports/daily-v2/2026-08-14.html |
| Grok 4.20 appears in API speed tracking | ๐ก partial | โ ๏ธ sensitive | 537ms avg total; 484ms avg TTFT | https://www.ailatency.com/reports/daily-v2/2026-08-03.html |
| Grok 4.3 appears in a later August speed snapshot | ๐ก partial | โ ๏ธ sensitive | 1.7s avg total; 698ms avg TTFT | https://www.ailatency.com/reports/daily-v2/2026-08-14.html |
| August API latency/cost tracking reported an overall Grade C and weighted-average latency around 1.6s | ๐ข confirmed | โ ๏ธ sensitive | Grade C; 1.6s | https://www.ailatency.com/reports/daily-v2/2026-08-03.html |
| Firecrawl relaunched a free keyless agent web search product | ๐ข confirmed | โ ๏ธ sensitive | free; sub-3-second | https://explainx.ai/catch-up-on-ai/2026-08-29 |
| GLM-5.3 open weights landed on Hugging Face on Aug. 28โ29, 2026 after an API-first launch | ๐ข confirmed | ๐ข robust | 2026-08-28/29; 753B total | https://chinaaibench.com/news/2026-08-31/ |
| GLM-5.3-Flash is described as 320B total / 18B active with 1M context | ๐ก partial | โ ๏ธ sensitive | 320B / 18B / 1M | https://www.requesty.ai/blog/open-weight-frontier-august-2026-glm-qwen-hy4 |
| Qwen3.8-Flash is described with API pricing around $0.16 / $0.47 per million tokens | ๐ก partial | โ ๏ธ sensitive | $0.16 / $0.47 | https://www.requesty.ai/blog/open-weight-frontier-august-2026-glm-qwen-hy4 |
| Hy4 Preview enters the frontier with very large scale | ๐ข confirmed | โ ๏ธ sensitive | 770B total / 49B active; 1M context | https://www.requesty.ai/blog/open-weight-frontier-august-2026-glm-qwen-hy4 |
| MiniMax M3 is described as a 1M-context multimodal open-weight model | ๐ข confirmed | โ ๏ธ sensitive | 428B open-weight MoE | https://www.requesty.ai/blog/open-weight-frontier-august-2026-glm-qwen-qwen-hy4 |
| DeepSeek V4-Flash-Vision-Exp is framed around multimodal agent benchmarks | ๐ก partial | โ ๏ธ sensitive | no single number in the claim | https://www.requesty.ai/blog/open-weight-frontier-august-2026-glm-qwen-hy4 |
Only two dual-pass matches held the same label (Claude Sonnet 5 price; GLM-5.3 open-weight landing). GLM-5.3-Flash specs and Qwen3.8-Flash price flipped ๐กโ๐ข and stay โ ๏ธ sensitive.