Ledger on GitHub. Hub here. Branches elsewhere.
drop-2026-09-13-kr-am
| claim | label | order | numbers | URL |
|---|---|---|---|---|
| Kimi K3 hit #1 on Arena.ai’s Frontend Code leaderboard, ahead of Claude Fable 5 and GPT-5.6 Sol. | 🟢 robust | robust | 1679 vs 1631 vs 1618 | https://kimik2ai.com/k3/ |
| Kimi K3 is described as a 2.8T-parameter open-weight model. | 🟢 robust | robust | 2.8T | https://kimik2ai.com/k3/ |
| DeepSeek V4-Pro was described as a 1.7T-class MoE with open weights and a low per-token price. | ⚠️ sensitive | partial | 1.7T; 0.435/0.87 per million tokens | https://mungomash.com/ai/models/ |
| Alibaba Qwen3.8-Max was reported as a 2.4T frontier model with a 1M-token context window. | 🟢 robust | robust | 2.4T; 1M tokens | https://www.forbes.com/sites/tylerroush/2026/08/03/alibaba-unveils-qwen38-max-model-chinas-latest-ai-challenger-to-openai-and-anthropic/ |
| OpenAI’s InferenceX comparison reportedly showed better AI work per watt than Nvidia Blackwell systems. | ⚠️ sensitive | partial | 1.5–1.9× | https://chinaaibench.com/news/2026-08-26/ |
| In the same comparison, DeepSeek R1 and Kimi K2.5 were reported at much higher per-user throughput than Blackwell systems. | ⚠️ sensitive | partial | 700 vs 169; 694 vs 182 tokens/s | https://chinaaibench.com/news/2026-08-26/ |
| Chinese open-weight models were said to be closing the gap with top closed models on cyber/bio capability. | ⚠️ sensitive | hype | no stable number | https://techcrunch.com/2026/08/04/open-weight-ai-models-are-catching-up-to-the-frontier-the-safety-gap-remains/ |
| Kimi K3 was reported to rank fourth on the Artificial Analysis Intelligence Index. | ⚠️ sensitive | partial | 4th | https://www.forbes.com/sites/drewbernstein/2026/08/03/chinese-ai-models-at-the-frontier/ |
| DeepSeek V4-Flash was described as a faster, cheaper tier with an AA Intelligence Index near GPT-5.6 Luna. | ⚠️ sensitive | hype | 50 vs 51; ~60% lower task cost | https://stochasticsandbox.com/posts/llm-encyclopedia-2026-08-08/ |
| Kimi K3 was released in late July 2026 as a 2.8T-parameter open-weight frontier model and hit the Arena.ai Frontend Code leaderboard. | 🟢 robust | robust | 2.8T; 104B active; 1679 vs 1631 | https://kimik2ai.com/k3/ |
| DeepSeek V4-Flash-0731 shipped under MIT/open-weight licensing with low API pricing and a 1M-token context window. | 🟢 robust | robust | 304B total; 13B active; ~$0.14/$0.28 per million tokens; 1M context | https://www.together.ai/models/deepseek-v4-flash-0731 |
| Qwen3.8-Max / Qwen3.8-2.4T-A95B was reported as a 2.4T frontier model with 95B active parameters and 1M-token context. | 🟢 robust | robust | 2.4T total; 95B active; 1M context | https://www.forbes.com/sites/tylerroush/2026/08/03/alibaba-unveils-qwen38-max-model-chinas-latest-ai-challenger-to-openai-and-anthropic/ |
| GLM-5.3-Flash was reported as a 321B/18B MoE open-weight model under MIT with strong coding benchmark scores. | ⚠️ sensitive | partial | 321B total; 18B active; 84.3 Terminal Bench 2.1; 63.4 DeepSWE 1.1 | https://forkast.news/chinese-open-weight-frontier-compresses-five-labs-thirty-days-two-licensing-models/ |
| Chinese open-weight frontier releases compressed into roughly a 30-day window across four flagship models, with two MIT and two custom-revenue-gated licenses. | 🟢 robust | robust | 30 days; 4 models; 2 MIT; 2 custom | https://forkast.news/chinese-open-weight-frontier-compresses-five-labs-thirty-days-two-licensing-models/ |
| Tencent’s Hy4-preview was a 770B-parameter MoE model with a 1M-token context window and Apache 2.0 release. | � |