Ledger on GitHub. Hub here. Branches elsewhere.
drop-2026-09-10-us-am
| claim | label | order | numbers | URL |
|---|---|---|---|---|
| Kimi K3 topped Arena’s frontend coding leaderboard, ahead of Claude Fable 5. | 🟢 robust | Kimi K3 #1 vs Fable 5 #2 | 1,679 vs 1,631 | https://testingmodels.com/news/kimi-k3-tops-frontend-code-arena |
| Kimi K3 beat Fable 5 and GPT-5.6 Sol on Arena’s frontend code leaderboard. | 🟢 robust | Kimi K3 #1; Fable 5 #2; GPT-5.6 Sol #3 | 1,679; 1,631; 1,618 | https://www.notebookcheck.net/Kimi-K3-tops-AI-benchmark-in-a-first-for-Chinese-models.1347112.0.html |
| DeepSeek V4-Pro-0813 scored 87.9 on Terminal-Bench 2.1, up from 72.1 in the preview build. | 🟢 robust | +15.8 | 87.9 vs 72.1 | https://datanorth.ai/news/deepseek-releases-v4-pro-0813-and-harness-v0-1 |
| DeepSeek V4-Pro-0813 also posted 62.7 on DeepSWE and 74.1 on Toolathlon-Verified. | 🟢 robust | DeepSWE; Toolathlon-Verified | 62.7; 74.1 | https://build.nvidia.com/deepseek-ai/deepseek-v4-pro-0813/modelcard |
| Independent coverage says DeepSeek V4 Pro’s strongest public number is vendor-reported, with a cooler independent picture. | ⚠️ sensitive | neutral harness; AI index | 79%; 53 | https://zhuermu.com/en/blog/deepseek-v4-pro-0813/ |
| Alibaba’s Qwen3.8-Max was reported as a 2.4T MoE with 95B active parameters, plus a promised 27B open-weight follow-on. | ⚠️ sensitive | 2.4T; 95B; 27B | https://kafkai.ai/articles/ai-technology/ai-models-july-august-2026-overview/ | |
| Chinese open-weight models reportedly reached 62% share on Vercel’s AI Gateway by late August, with DeepSeek-V4-Flash the most-used model. | ⚠️ sensitive | 62% | https://kafkai.ai/articles/ai-technology/ai-models-july-august-2026-overview/ | |
| Z.ai’s GLM-5.3 family was reported at 753B total parameters, 40B active parameters, and 1M context. | ⚠️ sensitive | 753B; 40B; 1M | https://fruition.net/frontier/ | |
| Tencent’s Hy4-preview was reported with weights and competitive coding scores. | ⚠️ sensitive | not specified | https://fruition.net/frontier/ | |
| Huawei’s Atlas 950 SuperPoD was presented as an 8,192-chip supernode with a 524 EFLOPS FP8 claim. | ⚠️ sensitive | 8,192; 524 EFLOPS FP8 | https://www.ainchina.com/blog/huawei-atlas-950-superpod-china-ai-chip-independence-2026/ | |
| OpenAI reportedly paused some frontier RL training to meet alignment, security, and monitoring standards. | ⚠️ sensitive | 2 weeks | https://thehackernews.com/2026/08/openai-pauses-frontier-rl-training-as.html |