Ledger on GitHub. Hub here. Branches elsewhere.
Eight labs cross 50 on the Intelligence Index in August 2026
Eight labs cross 50 on the Intelligence Index in August 2026
Alibaba shipped Qwen3.8-Max in early August 2026 as a sparse mixture-of-experts flagship with 2.4 trillion total parameters, 95 billion active per forward pass, and one million token context window.
The weights were published mid-August under open licenses. The model was described as Alibaba’s largest and most capable flagship to date.
DeepSeek shipped V4-Pro-0813 with 1.7 trillion parameters and MIT-licensed weights in the same window. The release was characterized as benchmark-parity with Opus 4.8 on agentic coding tasks, though no specific score was provided for that comparison.
DeepSWE results for V4-Pro-0813 reached 62.7, a gain of 49.9 points over the earlier preview snapshot.
Zhipu GLM-5.3 was documented as a 753 billion total parameter model with 40 billion active and one million token context. No public weight release was confirmed in the sources.
A serving benchmark compared GLM 5.3 on Nvidia hardware via SGLang against AMD, claiming up to five times better cost efficiency on Nvidia at 150 output tokens per second per user. The benchmark was attributed to Nvidia-backed testing published in a Forbes piece in late August.
Muse Glimmer 30B was described as an open-weight agentic model with a 2 billion parameter vision encoder and 28 billion parameter text decoder. The model was positioned as a Meta release from the Silicon Valley side.
SWE-Bench Verified performance for Muse Glimmer 30B was reported at 76.0.
The frontier tracker claimed that eight labs crossed the 50-point threshold on the Intelligence Index in August 2026. No lab names or specific scores were listed in the table.
All four major releases—Qwen3.8-Max, DeepSeek V4-Pro-0813, GLM-5.3, and Muse Glimmer 30B—were sparse MoE or multi-modal architectures. All carried context windows at or near one million tokens.
Three of the four published open weights under permissive licenses within days or weeks of announcement. GLM-5.3 was the exception; no weight publication was mentioned.
The August window produced a cluster of frontier-adjacent releases from China-based labs and one from Meta. DeepSeek’s MIT license and Alibaba’s open-weight commitment marked a continuation of the open-weights trajectory that began in mid-2025.
The Nvidia versus AMD serving comparison was narrow in scope: one model, one serving engine, one hardware generation. The five-times cost efficiency claim applied only to that specific stack and workload profile.
SWE-Bench Verified and DeepSWE scores appeared for two models. Muse Glimmer 30B scored 76.0 on SWE-Bench Verified. DeepSeek V4-Pro-0813 scored 62.7 on DeepSWE with a 49.9-point delta over preview.
No MMLU, GPQA, or other academic benchmark numbers appeared in the table. The releases were characterized by scale, sparsity, context length, and agentic task performance rather than zero-shot question-answering.
The Intelligence Index threshold of 50 was crossed by eight labs in the same month. The index itself was not defined in the sources, and no lab names or model names were tied to specific scores above 50.
Summary table
| Model | Lab | Parameters | Active | Context | License | Benchmark | Score | URL |
|---|---|---|---|---|---|---|---|---|
| Qwen3.8-Max | Alibaba | 2.4T | 95B | 1M | Open (mid-Aug) | — | — | link |
| DeepSeek V4-Pro-0813 | DeepSeek | 1.7T | — | — | MIT | DeepSWE | 62.7 | link |
| GLM-5.3 | Zhipu | 753B | 40B | 1M | — | — | — | link |
| Muse Glimmer 30B | Meta | 30B | — | — | Open | SWE-Bench Verified | 76.0 | link |
Act: Pull Qwen3.8-Max or DeepSeek V4-Pro-0813 weights if your stack can route across 95B or 1.7T sparse parameters and your tasks match the agentic coding or long-context profile.
Watch: The Intelligence Index threshold and the eight labs above 50; no lab names or scores were public in August, so the next disclosure will show whether the index correlates with SWE-Bench, DeepSWE, or internal routing metrics.
Ignore: The Nvidia-versus-AMD serving benchmark unless you run GLM 5.3 on SGLang at exactly 150 output tokens per second per user; the five-times cost claim was narrow and hardware-specific.