dropkit.contents

Ledger on GitHub. Hub here. Branches elsewhere.

← Drops · 2026-09-15

drop-2026-09-15-kr-am

claimlabelordernumbersURL
Qwen3.8-Max benchmark set shows it near frontier agentic models on terminal/code tasks🟢 robustearlier pass confirms; later pass still confirmsTerminal Bench 2.1 86.6; SWE-bench Pro 67.7; PaperBench 93.0; GPQA Diamond 92.6; OSWorld-Verified 86.1https://www.developersdigest.tech/blog/qwen-3-8-max-release-2026
Qwen3.8-Max is a 2.4T-parameter MoE with 95B active parameters🟢 robustearlier pass confirms; later pass still confirms2.4T total; 95B activehttps://www.digitalapplied.com/blog/qwen3-8-max-full-release-benchmarks-open-weights
OpenAI Jalapeño was benchmarked against Nvidia Blackwell on public inference metrics🟢 robustearlier pass confirms; later pass still confirmsup to 1.9× more AI work per watt; up to 3.6× lower latencyhttps://www.forbes.com/sites/jonmarkmarkman/2026/08/27/openai-publishes-first-jalapeo-benchmarks-against-nvidia-blackwell/
DeepSeek V4 Pro posted high software-engineering benchmark results and pricing⚠️ sensitiveearlier pass partial; later pass partialSWE-bench Verified 80.6; LiveCodeBench 93.5; $0.435 / $0.87 per million tokenshttps://www.schym.de/projects/ai-hyperscaler-radar/2026/august
DeepSeek V4 Pro 0813 improved on DeepSWE and Terminal Bench⚠️ sensitiveearlier pass partial; later pass partialDeepSWE 62.7; Terminal Bench 2.1 87.9https://thursdai.news/releases/2026-08
DeepSeek V4-Flash-Vision-Exp was described as a 305B multimodal MoE with MIT weights⚠️ sensitiveearlier pass partial; later pass partial305Bhttp://ai-tldr.dev/releases/model/
Chinese open weights dominated 2026 releases above 20B parameters, with Apache 2.0 and MIT most common🟢 robustearlier pass confirms; later pass confirms178 releases; 59% Apache 2.0; 22% MIThttps://www.reuters.com/world/china/open-source-ai-models-china-2026-08-14/
Kimi K3 briefly led the Frontend Code Arena leaderboard⚠️ sensitiveearlier pass partial; later pass partial1,679 vs 1,631https://www.forbes.com/sites/jonmarkmarkman/2026/08/07/
Kimi K3 also ranked high on Artificial Analysis Intelligence Index⚠️ sensitiveearlier pass partial; later pass partial4th overallhttps://www.forbes.com/sites/jonmarkmarkman/2026/08/07/
Meta Muse Glimmer was described as a 30B open-weight agentic model runnable on a single 24GB GPU⚠️ sensitiveearlier pass partial; later pass partial30B; 24GB GPUhttps://thursdai.news/releases/2026-08