Ledger on GitHub. Hub here. Branches elsewhere.
drop-2026-09-15-kr-am
| claim | label | order | numbers | URL |
|---|---|---|---|---|
| Qwen3.8-Max benchmark set shows it near frontier agentic models on terminal/code tasks | 🟢 robust | earlier pass confirms; later pass still confirms | Terminal Bench 2.1 86.6; SWE-bench Pro 67.7; PaperBench 93.0; GPQA Diamond 92.6; OSWorld-Verified 86.1 | https://www.developersdigest.tech/blog/qwen-3-8-max-release-2026 |
| Qwen3.8-Max is a 2.4T-parameter MoE with 95B active parameters | 🟢 robust | earlier pass confirms; later pass still confirms | 2.4T total; 95B active | https://www.digitalapplied.com/blog/qwen3-8-max-full-release-benchmarks-open-weights |
| OpenAI Jalapeño was benchmarked against Nvidia Blackwell on public inference metrics | 🟢 robust | earlier pass confirms; later pass still confirms | up to 1.9× more AI work per watt; up to 3.6× lower latency | https://www.forbes.com/sites/jonmarkmarkman/2026/08/27/openai-publishes-first-jalapeo-benchmarks-against-nvidia-blackwell/ |
| DeepSeek V4 Pro posted high software-engineering benchmark results and pricing | ⚠️ sensitive | earlier pass partial; later pass partial | SWE-bench Verified 80.6; LiveCodeBench 93.5; $0.435 / $0.87 per million tokens | https://www.schym.de/projects/ai-hyperscaler-radar/2026/august |
| DeepSeek V4 Pro 0813 improved on DeepSWE and Terminal Bench | ⚠️ sensitive | earlier pass partial; later pass partial | DeepSWE 62.7; Terminal Bench 2.1 87.9 | https://thursdai.news/releases/2026-08 |
| DeepSeek V4-Flash-Vision-Exp was described as a 305B multimodal MoE with MIT weights | ⚠️ sensitive | earlier pass partial; later pass partial | 305B | http://ai-tldr.dev/releases/model/ |
| Chinese open weights dominated 2026 releases above 20B parameters, with Apache 2.0 and MIT most common | 🟢 robust | earlier pass confirms; later pass confirms | 178 releases; 59% Apache 2.0; 22% MIT | https://www.reuters.com/world/china/open-source-ai-models-china-2026-08-14/ |
| Kimi K3 briefly led the Frontend Code Arena leaderboard | ⚠️ sensitive | earlier pass partial; later pass partial | 1,679 vs 1,631 | https://www.forbes.com/sites/jonmarkmarkman/2026/08/07/ |
| Kimi K3 also ranked high on Artificial Analysis Intelligence Index | ⚠️ sensitive | earlier pass partial; later pass partial | 4th overall | https://www.forbes.com/sites/jonmarkmarkman/2026/08/07/ |
| Meta Muse Glimmer was described as a 30B open-weight agentic model runnable on a single 24GB GPU | ⚠️ sensitive | earlier pass partial; later pass partial | 30B; 24GB GPU | https://thursdai.news/releases/2026-08 |