August 2026: harness, open weights, and the agent stack
DeepSeek Harness, GLM-5.3, Qwen3.8-Max, and a homelab RTX 3060 datapoint.
Ledger on GitHub. Hub here. Branches elsewhere.
DeepSeek Harness, GLM-5.3, Qwen3.8-Max, and a homelab RTX 3060 datapoint.
Two open-weight MoE models crossed 700B total parameters this cycle. One agentic benchmark crossed 80% on SWE-bench Verified. One Apache 2.0 open-weight model at 30B reached 76.0 o
AI firms still lack sufficient containment measures; no lab reached full implementation.
A latency snapshot for Grok 4.20 reports 538 ms total and 513 ms TTFT. Gemini 2.5 Flash latency is 667 ms total and 506 ms TTFT, with an output cost of $0.0025 per 1K output tokens
This publication briefing summarizes observations from the field. It covers recent shifts in platform positioning and open-weight model statistics.
Gemini, Claude, Grok pricing and a one-page scoreboard from public sources.
2026-08-29 lab note: Kimi K3 / Qwen 2.4T MoE vs US open weights; Gemini, Claude, ChatGPT, Grok.