Ledger on GitHub. Hub here. Branches elsewhere.
drop-2026-09-12-kr-pm
| claim | label | order (🟢 robust / ⚠️ sensitive) | numbers | URL |
|---|---|---|---|---|
| Grok 4.6 led the newcomers in Agents on Rails, finishing 52 of 63 runs and posting 33% Rails API recall. | confirmed | 🟢 robust | 52/63; 33% | https://rubyonrails.org/2026/8/17/agents-on-rails-grok-4-6-glm-5-3-gemini-3-7-flash-and-opus-4-8 |
| Claude Opus 4.8 ranked ahead of Grok 4.6 on Rails API recall, at 35% versus 33%. | confirmed | 🟢 robust | 35% vs 33% | https://rubyonrails.org/2026/8/17/agents-on-rails-grok-4-6-glm-5-3-gemini-3-7-flash-and-opus-4-8 |
| Gemini 3.7 Flash sat mid-pack in Agents on Rails with 45 of 63 runs completed and a lower campaign cost than Grok 4.6. | confirmed | 🟢 robust | 45/63; $17.85 | https://rubyonrails.org/2026/8/17/agents-on-rails-grok-4-6-glm-5-3-gemini-3-7-flash-and-opus-4-8 |
| Grok 4.6 pricing was reported around $2.00/M input, $0.50/M cached input, and $6.00/M output. | partial | ⚠️ sensitive | $2.00; $0.50; $6.00 | https://www.aipricing.guru/news/xai-grok-4-6-launch-pricing-impact-august-2026/ |
| Grok 4.6 was also described as having about 500K-token context in a comparison page. | partial | ⚠️ sensitive | 500K; $0.50/M | https://rohitai.com/blog/best-ai-models-2026-openai-anthropic-google-xai-deepseek |
| GLM-5.3-Flash was described in roundup coverage as 320B total, 18B active, with 1M context and MIT licensing. | partial | ⚠️ sensitive | 320B; 18B; 1M | https://www.requesty.ai/blog/open-weight-frontier-august-2026-glm-qwen-hy4 |
| GLM-5.3 was reported as competitive with Claude on agentic tasks in public coverage. | partial | ⚠️ sensitive | n/a | https://gigazine.net/gsc_news/en/20260829-glm-5-3-open/ |
| GLM-5.3 safety coverage said it lagged GPT-5.5 and Claude Opus 4.7 on cyber/bio capability, while showing no refusals on offensive tasks in that test set. | partial | ⚠️ sensitive | n/a | https://techcrunch.com/2026/08/04/open-weight-ai-models-are-catching-up-to-the-frontier-the-safety-gap-remains/ |
| Qwen3.8-Flash was described as a multimodal MoE preview with 125B parameters plus a 51B n-gram component. | partial | ⚠️ sensitive | 125B; 51B | https://www.requesty.ai/blog/open-weight-frontier-august-2026-glm-qwen-hy4 |
| Qwen3.8-Flash API pricing was listed around $0.16/M input and $0.47/M output. | partial | ⚠️ sensitive | $0.16; $0.47 | https://www.requesty.ai/blog/open-weight-frontier-august-2026-glm-qwen-hy4 |
| Qwen3.8-Max was presented as a frontier model with about 2.4T parameters and 1M context. | partial | ⚠️ sensitive | 2.4T; 1M | https://llm.okamomedia.tokyo/en/china-ai/ |
| Kimi K3 was said to have full weights with 2.8T parameters and 104B active parameters. | partial | ⚠️ sensitive | 2.8T; 104B | https://www.requesty.ai/blog/open-weight-frontier-august-2026-glm-qwen-hy4 |
Directly comparable “Gemini vs Claude vs ChatGPT vs Grok” single-source coverage was not fully confirmed in the gathered URLs, so the table keeps only source-backed claims and downgrades secondary-angle items.