dropkit.contents

Ledger on GitHub. Hub here. Branches elsewhere.

← Drops · 2026-09-08

drop-2026-09-08-us-am

claimlabelordernumbersURL
OpenAI’s Jalapeño chip shows internal inference gains vs commercial NVIDIA systems🟢 robustrobust1.5–1.9× more AI work per watt; 1.7–3.6× lower end-to-end latency; 2.1–4.1× higher performance on interactive workloads[1][2]https://openai.com/index/jalapeno-first-results/
OpenAI and Broadcom co-designed Jalapeño for inference, with engineering samples running production-target workloads🟢 robustrobustfirst custom inference chip; engineering samples; production-target workloads[2]https://openai.com/index/the-full-stack-behind-abundant-intelligence/
NVIDIA says Qwen3.8-2.4T-A95B hits day-0 serve throughput on GB300 NVL72🟢 robustrobust2.4T total / 95B active; >4,000 tok/s/GPU; >350 tok/s/user; 72 GPUs; 130 TB/s NVLink[3][4][5]https://developer.nvidia.com/blog/serve-qwen3-8-2-4t-a95b-a-2-4t-parameter-model-with-configurable-reasoning-on-nvidia-gb300-nvl72/
GLM-5.3 benchmark jump on Terminal-Bench 3.0 is large but still below some closed models🟢 robustrobust4.6 → 28.3; 23.8 → 28.5; 77.2 → 84.5; TB3.0 still below 33.7 and 34.6 in cited comparison[6][7]https://huggingface.co/zai-org/GLM-5.3
Morph and related coverage repeat the same GLM-5.3 benchmark deltas⚠️ sensitivesensitive4.6 → 28.3[8]https://www.morphllm.com/glm-5-3
Zhipu/GLM-5.3 still trails top closed models on Terminal-Bench 3.0⚠️ sensitivesensitive28.3 vs 33.7 vs 34.6[9]https://zenn.dev/neotechpark/articles/7003e20300d7e5
DeepSeek-V4-Flash-0731 is presented as a smaller MoE with strong independent benchmarking⚠️ sensitivesensitive284B total / 13B active; 1M in / 384k out; AA 50; TB 82.7%[10]https://www.deeplearning.ai/the-batch/deepseek-pushes-the-frontier-again-deepseek-refreshed-its-v4-flash-model-with-an-impressive-fine-tune
DeepSeek founder leaked-call claims about compute scarcity and training scale are highly specific but source quality is lower⚠️ sensitivesensitive~20,000 H-equivalent GPUs; 50k GB300s or 200k Huawei 950s; ¥50bn; 1–2 years behind[11]https://dealroom.co/news/143742-deepseeks-leaked-investor-call-1-20th-the-compute-agi-or-nothing-and-the-
ByteDance reportedly trained a massive model to rival Anthropic⚠️ sensitivesensitiveno reliable numbers in gathered coverage[12]https://arstechnica.com/ai/2026/08/bytedance-trains-massive-ai-model-in-bid-to-rival-anthropic/
Moore Threads’ same-day GLM-5.3-Flash support on MUSA/SGLang-MUSA shows domestic chip/inference integration⚠️ sensitivesensitivesame-day support; MTT S5000[13]https://www.techtimes.com/articles/325872/20260828/sanctioned-chinese-chips-just-served-62-trillion-ai-tokens-frontier-scale.htm

[1] https://openai.com/index/jalapeno-first-results/
[2] https://openai.com/index/the-full-stack-behind-abundant-intelligence/
[3] https://developer.nvidia.com/blog/serve-qwen3-8-2-4t-a95b-a-2-4t-parameter-model-with-configurable-reasoning-on-nvidia-gb300-nvl72/
[4] https://developer.nvidia.com/blog/experiment-with-qwen3-8-flash-next-on-nvidia-gb300-nvl72-for-agentic-coding/
[5] https://www.eenewseurope.com/en/nvidia-qwen3-8-model-gb300-nvl72/
[6] https://huggingface.co/zai-org/GLM-5.3
[7] https://huggingface.co/ssji3554/GLM-5.3
[8]