Ledger on GitHub. Hub here. Branches elsewhere.
drop-2026-09-12-us-am
| claim | label | order | numbers | URL |
|---|---|---|---|---|
| Qwen3.8-Max shipped as a sparse MoE flagship with 2.4T total parameters and about 95B active per step. | 🟢 robust | 1 | 2.4T total; 95B active | https://docs.qwencloud.com/changelog/models |
| Qwen3.8-Max was released in early August 2026 with 1M-token context and $2/$6 per 1M-token pricing. | 🟢 robust | 2 | 1M context; $2/$6 per 1M tokens | https://docs.qwencloud.com/changelog/models |
| Alibaba published the open-weight base Qwen3.8-2.4T-A95B after the API launch. | 🟢 robust | 3 | 2.4T total; 95B active | https://rits.shanghai.nyu.edu/ai/qwen3-8-2-4t-a95b-alibaba-open-weights-its-max-tier-flagship |
| Qwen3.8-27B was reported as a dense open-weight model positioned to match stronger Qwen variants on consumer hardware. | ⚠️ sensitive | 4 | 27B | https://www.developersdigest.tech/blog/qwen-3-8-max-release-2026 |
| DeepSeek-V4-Flash was reported with production-style inference metrics around 2,400 tok/s on a single A100, plus 32 ms p50 first-token latency and 118 ms p99. | ⚠️ sensitive | 5 | 2,400 tok/s; 32 ms p50; 118 ms p99 | https://mr.technology/payloads/i-ran-eleven-open-weights-inference-stacks-production-agent-traffic-august-2026 |
| Open-weight routing data showed Chinese models taking nearly half of OpenRouter tokens. | ⚠️ sensitive | 6 | nearly 50% | https://www.forbes.com/sites/drewbernstein/2026/08/03/chinese-ai-models-at-the-frontier/ |
| Chinese open-weight models reached a record 62% share on Vercel’s AI Gateway, and DeepSeek-V4-Flash became the most-used model there. | 🟢 robust | 7 | 62% | https://kafkai.ai/articles/ai-technology/ai-models-july-august-2026-overview/ |
| Tencent Hy4 Preview was described as a 770B total / 49B active MoE with a 1M-token context window. | ⚠️ sensitive | 8 | 770B; 49B; 1M | https://buttondown.com/patricknovak1/archive/models-agents-weekly-issue-018/ |
| GLM-5.3-Flash was described as a 320B-parameter MoE with 18B active per token, aimed at low-cost inference. | ⚠️ sensitive | 9 | 320B; 18B | https://www.techtimes.com/articles/325858/20260828/ox-alpha-was-glm-53-flash-china-inference-chip-claim-stands-unverified.htm |
| NVIDIA was reported to be planning a China-focused inference chip using licensed Groq LPU technology alongside GPUs. | ⚠️ sensitive | 10 | — | https://www.theneuron.ai/digest/everything-that-happened-in-ai-today-thursday-august-20-2026/ |
| OpenAI’s Jalapeño inference chip was reported with 1.5-1.9x more work per watt and 1.7-3.6x lower end-to-end latency versus prior systems. | ⚠️ sensitive | 11 | 1.5-1.9x; 1.7-3.6x | https://paragraph.com/@twiata/this-week-in-all-things-ai-week-35-2026 |