dropkit.contents

Ledger on GitHub. Hub here. Branches elsewhere.

← Drops · 2026-09-12

drop-2026-09-12-us-am

claimlabelordernumbersURL
Qwen3.8-Max shipped as a sparse MoE flagship with 2.4T total parameters and about 95B active per step.🟢 robust12.4T total; 95B activehttps://docs.qwencloud.com/changelog/models
Qwen3.8-Max was released in early August 2026 with 1M-token context and $2/$6 per 1M-token pricing.🟢 robust21M context; $2/$6 per 1M tokenshttps://docs.qwencloud.com/changelog/models
Alibaba published the open-weight base Qwen3.8-2.4T-A95B after the API launch.🟢 robust32.4T total; 95B activehttps://rits.shanghai.nyu.edu/ai/qwen3-8-2-4t-a95b-alibaba-open-weights-its-max-tier-flagship
Qwen3.8-27B was reported as a dense open-weight model positioned to match stronger Qwen variants on consumer hardware.⚠️ sensitive427Bhttps://www.developersdigest.tech/blog/qwen-3-8-max-release-2026
DeepSeek-V4-Flash was reported with production-style inference metrics around 2,400 tok/s on a single A100, plus 32 ms p50 first-token latency and 118 ms p99.⚠️ sensitive52,400 tok/s; 32 ms p50; 118 ms p99https://mr.technology/payloads/i-ran-eleven-open-weights-inference-stacks-production-agent-traffic-august-2026
Open-weight routing data showed Chinese models taking nearly half of OpenRouter tokens.⚠️ sensitive6nearly 50%https://www.forbes.com/sites/drewbernstein/2026/08/03/chinese-ai-models-at-the-frontier/
Chinese open-weight models reached a record 62% share on Vercel’s AI Gateway, and DeepSeek-V4-Flash became the most-used model there.🟢 robust762%https://kafkai.ai/articles/ai-technology/ai-models-july-august-2026-overview/
Tencent Hy4 Preview was described as a 770B total / 49B active MoE with a 1M-token context window.⚠️ sensitive8770B; 49B; 1Mhttps://buttondown.com/patricknovak1/archive/models-agents-weekly-issue-018/
GLM-5.3-Flash was described as a 320B-parameter MoE with 18B active per token, aimed at low-cost inference.⚠️ sensitive9320B; 18Bhttps://www.techtimes.com/articles/325858/20260828/ox-alpha-was-glm-53-flash-china-inference-chip-claim-stands-unverified.htm
NVIDIA was reported to be planning a China-focused inference chip using licensed Groq LPU technology alongside GPUs.⚠️ sensitive10https://www.theneuron.ai/digest/everything-that-happened-in-ai-today-thursday-august-20-2026/
OpenAI’s Jalapeño inference chip was reported with 1.5-1.9x more work per watt and 1.7-3.6x lower end-to-end latency versus prior systems.⚠️ sensitive111.5-1.9x; 1.7-3.6xhttps://paragraph.com/@twiata/this-week-in-all-things-ai-week-35-2026