Ledger on GitHub. Hub here. Branches elsewhere.
drop-2026-09-08-us-am
| claim | label | order | numbers | URL |
|---|---|---|---|---|
| OpenAI’s Jalapeño chip shows internal inference gains vs commercial NVIDIA systems | 🟢 robust | robust | 1.5–1.9× more AI work per watt; 1.7–3.6× lower end-to-end latency; 2.1–4.1× higher performance on interactive workloads[1][2] | https://openai.com/index/jalapeno-first-results/ |
| OpenAI and Broadcom co-designed Jalapeño for inference, with engineering samples running production-target workloads | 🟢 robust | robust | first custom inference chip; engineering samples; production-target workloads[2] | https://openai.com/index/the-full-stack-behind-abundant-intelligence/ |
| NVIDIA says Qwen3.8-2.4T-A95B hits day-0 serve throughput on GB300 NVL72 | 🟢 robust | robust | 2.4T total / 95B active; >4,000 tok/s/GPU; >350 tok/s/user; 72 GPUs; 130 TB/s NVLink[3][4][5] | https://developer.nvidia.com/blog/serve-qwen3-8-2-4t-a95b-a-2-4t-parameter-model-with-configurable-reasoning-on-nvidia-gb300-nvl72/ |
| GLM-5.3 benchmark jump on Terminal-Bench 3.0 is large but still below some closed models | 🟢 robust | robust | 4.6 → 28.3; 23.8 → 28.5; 77.2 → 84.5; TB3.0 still below 33.7 and 34.6 in cited comparison[6][7] | https://huggingface.co/zai-org/GLM-5.3 |
| Morph and related coverage repeat the same GLM-5.3 benchmark deltas | ⚠️ sensitive | sensitive | 4.6 → 28.3[8] | https://www.morphllm.com/glm-5-3 |
| Zhipu/GLM-5.3 still trails top closed models on Terminal-Bench 3.0 | ⚠️ sensitive | sensitive | 28.3 vs 33.7 vs 34.6[9] | https://zenn.dev/neotechpark/articles/7003e20300d7e5 |
| DeepSeek-V4-Flash-0731 is presented as a smaller MoE with strong independent benchmarking | ⚠️ sensitive | sensitive | 284B total / 13B active; 1M in / 384k out; AA 50; TB 82.7%[10] | https://www.deeplearning.ai/the-batch/deepseek-pushes-the-frontier-again-deepseek-refreshed-its-v4-flash-model-with-an-impressive-fine-tune |
| DeepSeek founder leaked-call claims about compute scarcity and training scale are highly specific but source quality is lower | ⚠️ sensitive | sensitive | ~20,000 H-equivalent GPUs; 50k GB300s or 200k Huawei 950s; ¥50bn; 1–2 years behind[11] | https://dealroom.co/news/143742-deepseeks-leaked-investor-call-1-20th-the-compute-agi-or-nothing-and-the- |
| ByteDance reportedly trained a massive model to rival Anthropic | ⚠️ sensitive | sensitive | no reliable numbers in gathered coverage[12] | https://arstechnica.com/ai/2026/08/bytedance-trains-massive-ai-model-in-bid-to-rival-anthropic/ |
| Moore Threads’ same-day GLM-5.3-Flash support on MUSA/SGLang-MUSA shows domestic chip/inference integration | ⚠️ sensitive | sensitive | same-day support; MTT S5000[13] | https://www.techtimes.com/articles/325872/20260828/sanctioned-chinese-chips-just-served-62-trillion-ai-tokens-frontier-scale.htm |
[1] https://openai.com/index/jalapeno-first-results/
[2] https://openai.com/index/the-full-stack-behind-abundant-intelligence/
[3] https://developer.nvidia.com/blog/serve-qwen3-8-2-4t-a95b-a-2-4t-parameter-model-with-configurable-reasoning-on-nvidia-gb300-nvl72/
[4] https://developer.nvidia.com/blog/experiment-with-qwen3-8-flash-next-on-nvidia-gb300-nvl72-for-agentic-coding/
[5] https://www.eenewseurope.com/en/nvidia-qwen3-8-model-gb300-nvl72/
[6] https://huggingface.co/zai-org/GLM-5.3
[7] https://huggingface.co/ssji3554/GLM-5.3
[8]