dropkit.contents

Ledger on GitHub. Hub here. Branches elsewhere.

← Drops · 2026-09-12

drop-2026-09-12-kr-am

claimlabelordernumbersURL
Qwen3.8-2.4T-A95B is a text-only open-weights repo on Hugging Face that requires thinking mode and contains model weights/config files.🟢 confirmed🟢 robust2.4T total; 95B activehttps://huggingface.co/Qwen/Qwen3.8-2.4T-A95B
Qwen3.8-2.4T-A95B has a separate FP8 quantized companion checkpoint.🟢 confirmed🟢 robustFP8https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B-FP8
Qwen3.8-2.4T-A95B’s Hugging Face repo includes deployment examples using SGLang.🟢 confirmed🟢 robustSGLang; 30000 porthttps://huggingface.co/Qwen/Qwen3.8-2.4T-A95B/tree/main
Qwen3.8-2.4T-A95B is presented as a Max-class MoE release with 512 experts per layer and native 262K context, extensible to about 1M.🟡 partial⚠️ sensitive2.4T; ~95B active; 512 experts; 262K; ~1Mhttps://huggingface.co/Qwen/Qwen3.8-2.4T-A95B/discussions/27
Qwen3.8-2.4T-A95B has an FP8 deployment-focused repo that is compatible with vLLM, SGLang, and TokenSpeed.🟢 confirmed🟢 robustFP8; vLLM; SGLang; TokenSpeedhttps://huggingface.co/unsloth/Qwen3.8-2.4T-A95B-GGUF
GLM-5.3-Flash is reported as an open-weights MoE release with 320B total parameters, 18B active parameters, and a 1M-token context window.🟡 partial⚠️ sensitive320B; 18B; 1,048,576https://ai-tldr.dev/models/glm-5-3-flash/
GLM-5.3-Flash is reported to support native multimodal inputs and MIT-licensed open weights on Hugging Face.🟡 partial⚠️ sensitivemultimodal; MIThttps://ai-tldr.dev/models/glm-5-3-flash/
Z.ai’s GLM-5.3-Flash is described in secondary coverage as supporting SGLang, vLLM, TokenSpeed, Transformers, KTransformers, and Unsloth.🟡 partial⚠️ sensitive1M context; multiple serving stackshttps://ai-tldr.dev/models/glm-5-3-flash/
DeepSeek V4 Pro 0813 is reported by a vendor-facing API page as supporting a 1M-token context window.🟡 partial⚠️ sensitive1M contexthttps://developer.puter.com/ai/deepseek/deepseek-v4-pro-0813/
DeepSeek V4 Pro 0813 pricing is reported as $1.32 per 1M input tokens and $3.96 per 1M output tokens on one provider page, while other aggregators show different tariff bands.🟡 partial⚠️ sensitive$1.32 / $3.96 per 1Mhttps://developer.puter.com/ai/deepseek/deepseek-v4-pro-0813/
DeepSeek V4 Pro 0813 is also reported elsewhere with lower tariff-band figures, so its price should be treated as provider-dependent rather than fixed.🟡 partial⚠️ sensitive$0.435 / $0.87 per 1Mhttps://apidog.com/blog/how-to-use-deepseek-v4-pro-0813-api/
DeepSeek V4 Pro 0813 is listed by a third-party model tracker with a 1M context window and model-ranking metadata.🟡 partial⚠️ sensitive1M; rank #18; 52.37%https://www.vals.ai/models/deepseek_deepseek-v4-pro-0813
Claims that GLM-5.3-Flash ran on domestically produced Chinese AI chips with a custom SGLang-based stack are still indirect and not fully verified.🔴 hype⚠️ sensitiven/ahttps://www.techtimes.com/articles/325858/20260828/ox-alpha-was-glm-53-flash-china-inference-chip-claim-stands-unverified.htm
Claims that DeepSeek V4-Pro-0813 was “open-sourced” should be treated cautiously because the gathered evidence is inconsistent.🔴 hype⚠️ sensitiven/ahttps://www.datan