dropkit.contents

Ledger on GitHub. Hub here. Branches elsewhere.

← Drops · 2026-09-02

August 2026: harness, open weights, and the agent stack

August 2026: harness, open weights, and the agent stack

DeepSeek Harness shipped as an open-source MIT-licensed agent harness on August 13, 2026. The project entered developer preview with an “everything is a plugin” architecture.

The harness is built as a runtime for autonomous agents. Early adoption was strong: 28,000+ GitHub stars in approximately three hours, with some coverage citing 28,700. One digest compared it to Claude Code.

GLM-5.3 arrived with a 1M context window. The model uses 753B total parameters with 40B active. Some coverage noted the weights as staged or pending safety review rather than instantly public.

GLM-5.3 posted large benchmark gains on agent, coding, and cyber tasks versus the prior GLM-5.2 release. Terminal-Bench 3.0 rose from 4.6 to 28.3. DeepSWE v1.1 climbed from 46.2 to 66.9. CyberGym went from 77.2% to 84.5%. ExploitBench jumped from 24.4% to 54.4%.

One summary described GLM-5.3 as fully open, MIT-licensed, with 320B total parameters and 18B active. The 320B / 18B figure is distinct from the 753B / 40B figure in earlier sources; both are in the table. An independent ranking digest placed GLM-5.3 tied for 11th among open-weight models in August.

Alibaba announced Qwen3.8-Max as the most capable Qwen model to date. The model carries 2.4T parameters. Open weights were scheduled for release within the week following the announcement.

Secondary coverage described Qwen3.8-Max as frontier-tier, with strong agent and coding scores. Terminal-Bench 2.1 reached 86.6. OSWorld-Verified hit 86.1. SWE-bench Pro scored 67.7.

China’s open-model ecosystem was unusually large in 2026. The count of Chinese releases above 20B parameters reached 178. Of those, 59% carried Apache 2.0 licenses and 22% carried MIT licenses.

One report described Chinese frontier models as narrowing the gap with U.S. closed models on cyber and bio capability, characterizing the lag as “a few months behind.”

Meta’s Muse Glimmer was framed as a local-first agentic open-weights model aimed at consumer GPUs. The model is 30B parameters and Apache 2.0 licensed.


Scoreboard

The table below lists each claim with its confirmation label, order robustness, key numbers, and source URL.

| Claim | Label | Order | Numbers | URL | |---|---|---|---| | DeepSeek Harness is an open-source MIT-licensed agent harness released as a developer preview in August 2026. | confirmed | 🟢 robust | developer preview; MIT; launched 2026-08-13 | https://www.infoq.com/news/2026/08/deep-seek-harness/ | | DeepSeek Harness uses an “everything is a plugin” architecture and is built as a runtime for autonomous agents. | confirmed | 🟢 robust | plugin-based runtime; autonomous agents | https://github.com/deepseek-ai/deepseek-harness | | DeepSeek Harness was reported as a Claude Code-like open agent framework with strong early adoption. | partial | ⚠️ sensitive | 28,000+ GitHub stars in ~3 hours; 28,700 cited in coverage | https://temperature2.com/p/2026-08-13-deepseek-harness-open-source-agent-framework/ | | GLM-5.3 is an open-weight model with a 1M context window, but some coverage frames the weights as staged or pending safety review rather than instantly public. | partial | ⚠️ sensitive | 753B total / 40B active; 1M context | https://www.eigent.ai/blog/glm-5-3-coding-cyber-model | | GLM-5.3 posted large benchmark gains on agent, coding, and cyber tasks versus GLM-5.2. | confirmed | 🟢 robust | Terminal-Bench 3.0: 4.6→28.3; DeepSWE v1.1: 46.2→66.9; CyberGym: 77.2%→84.5%; ExploitBench: 24.4%→54.4% | https://www.eigent.ai/blog/glm-5-3-coding-cyber-model | | GLM-5.3 was summarized as a fully open MIT-licensed model with 320B total parameters and 18B active parameters. | confirmed | 🟢 robust | 320B total; 18B active; MIT | https://www.yottalabs.ai/post/glm-5-3-whats-new-benchmarks-how-to-access-it-2026 | | GLM-5.3 was independently indexed as one of the top open-weight models in an external ranking digest. | partial | ⚠️ sensitive | tied for 11th in one August digest | https://buttondown.com/theaggregate/archive/glm-53-enters-the-best-available-ranking-tied-for/ | | Qwen3.8-Max was announced as Alibaba’s most capable Qwen model, with open weights scheduled shortly after the announcement. | confirmed | 🟢 robust | 2.4T parameters; weights next week | https://www.alibabacloud.com/en/press-room/alibaba-unveils-qwen3-8-max?_p_lc=1 | | Qwen3.8-Max was described in secondary coverage as a frontier-tier open-weights model with strong agent and coding scores. | partial | ⚠️ sensitive | Terminal-Bench 2.1: 86.6; OSWorld-Verified: 86.1; SWE-bench Pro: 67.7 | https://www.marktechpost.com/2026/08/03/alibaba-qwen-releases-qwen3-8-max/ | | China’s open-model ecosystem was unusually large in 2026, with many releases above 20B parameters and a high share of permissive licenses. | confirmed | 🟢 robust | 178 Chinese releases above 20B; 59% Apache 2.0; 22% MIT | https://huggingface.co/blog/state-of-open-models-summer-2026 | | Chinese frontier models were reported to be narrowing the gap with U.S. closed models on cyber and bio capability. | partial | ⚠️ sensitive | “a few months behind” | https://techcrunch.com/2026/08/04/open-weight-ai-models-are-catching-up-to-the-frontier-the-safety-gap-remains/ | | Meta’s Muse Glimmer was framed as a local-first agentic open-weights model aimed at consumer GPUs. | partial | ⚠️ sensitive | 30B; Apache 2.0 | https://www.ideabosque.com/library/muse-glimmer-local-first-agentic-open-weights/ |


Practitioner lines

Act: Test DeepSeek Harness for autonomous agent workloads; the plugin architecture and MIT license lower integration cost. Benchmark GLM-5.3 and Qwen3.8-Max against your coding and agent eval suites; the Terminal-Bench and SWE-bench Pro numbers are high enough to justify eval cycles.

Watch: Track the staged-release pattern for open weights; the GLM-5.3 notes about safety review may signal a shift in how Chinese labs publish. Monitor the gap-narrowing claim on cyber and bio capability; “a few months behind” is imprecise but directional.

Ignore: Hype cycles around GitHub star velocity; 28,000 stars in three hours is notable but not a decision input without usage or integration data. Ranking ties (11th) in aggregated digests; treat as discovery signal, not capability proof.

Homelab box (not a vendor bench)

Same day, this factory’s RTX 3060 12GB ran Ollama qwen2.5:7b at 69.4 tok/s (64 eval tokens, 0.92s, ~4663 MiB VRAM, GPU 99%). Use only as our machine’s number.