Ledger on GitHub. Hub here. Branches elsewhere.
drop-2026-09-12-kr-am
| claim | label | order | numbers | URL |
|---|---|---|---|---|
| Qwen3.8-2.4T-A95B is a text-only open-weights repo on Hugging Face that requires thinking mode and contains model weights/config files. | 🟢 confirmed | 🟢 robust | 2.4T total; 95B active | https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B |
| Qwen3.8-2.4T-A95B has a separate FP8 quantized companion checkpoint. | 🟢 confirmed | 🟢 robust | FP8 | https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B-FP8 |
| Qwen3.8-2.4T-A95B’s Hugging Face repo includes deployment examples using SGLang. | 🟢 confirmed | 🟢 robust | SGLang; 30000 port | https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B/tree/main |
| Qwen3.8-2.4T-A95B is presented as a Max-class MoE release with 512 experts per layer and native 262K context, extensible to about 1M. | 🟡 partial | ⚠️ sensitive | 2.4T; ~95B active; 512 experts; 262K; ~1M | https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B/discussions/27 |
| Qwen3.8-2.4T-A95B has an FP8 deployment-focused repo that is compatible with vLLM, SGLang, and TokenSpeed. | 🟢 confirmed | 🟢 robust | FP8; vLLM; SGLang; TokenSpeed | https://huggingface.co/unsloth/Qwen3.8-2.4T-A95B-GGUF |
| GLM-5.3-Flash is reported as an open-weights MoE release with 320B total parameters, 18B active parameters, and a 1M-token context window. | 🟡 partial | ⚠️ sensitive | 320B; 18B; 1,048,576 | https://ai-tldr.dev/models/glm-5-3-flash/ |
| GLM-5.3-Flash is reported to support native multimodal inputs and MIT-licensed open weights on Hugging Face. | 🟡 partial | ⚠️ sensitive | multimodal; MIT | https://ai-tldr.dev/models/glm-5-3-flash/ |
| Z.ai’s GLM-5.3-Flash is described in secondary coverage as supporting SGLang, vLLM, TokenSpeed, Transformers, KTransformers, and Unsloth. | 🟡 partial | ⚠️ sensitive | 1M context; multiple serving stacks | https://ai-tldr.dev/models/glm-5-3-flash/ |
| DeepSeek V4 Pro 0813 is reported by a vendor-facing API page as supporting a 1M-token context window. | 🟡 partial | ⚠️ sensitive | 1M context | https://developer.puter.com/ai/deepseek/deepseek-v4-pro-0813/ |
| DeepSeek V4 Pro 0813 pricing is reported as $1.32 per 1M input tokens and $3.96 per 1M output tokens on one provider page, while other aggregators show different tariff bands. | 🟡 partial | ⚠️ sensitive | $1.32 / $3.96 per 1M | https://developer.puter.com/ai/deepseek/deepseek-v4-pro-0813/ |
| DeepSeek V4 Pro 0813 is also reported elsewhere with lower tariff-band figures, so its price should be treated as provider-dependent rather than fixed. | 🟡 partial | ⚠️ sensitive | $0.435 / $0.87 per 1M | https://apidog.com/blog/how-to-use-deepseek-v4-pro-0813-api/ |
| DeepSeek V4 Pro 0813 is listed by a third-party model tracker with a 1M context window and model-ranking metadata. | 🟡 partial | ⚠️ sensitive | 1M; rank #18; 52.37% | https://www.vals.ai/models/deepseek_deepseek-v4-pro-0813 |
| Claims that GLM-5.3-Flash ran on domestically produced Chinese AI chips with a custom SGLang-based stack are still indirect and not fully verified. | 🔴 hype | ⚠️ sensitive | n/a | https://www.techtimes.com/articles/325858/20260828/ox-alpha-was-glm-53-flash-china-inference-chip-claim-stands-unverified.htm |
| Claims that DeepSeek V4-Pro-0813 was “open-sourced” should be treated cautiously because the gathered evidence is inconsistent. | 🔴 hype | ⚠️ sensitive | n/a | https://www.datan |