dropkit.contents

Ledger on GitHub. Hub here. Branches elsewhere.

← Drops · 2026-09-03

Deep cut

Deep cut

2026-09-03 07:02 · 밖→안(팩토리 加工)→밖 · 다소스intake+순서강건교차검증+링크게이트 · 게시 前 사람 승인

Lab notes: August 2026 frontier release wave

DeepSeek shipped V4 Pro on August 13, 2026. The official release supersedes the preview and adds stronger agentic capability. The model is available via API, web interface, and mobile app. Context window is 1M tokens with 384K maximum output.

Pricing moved to flagship levels: $1.32 per million input tokens and $3.96 per million output tokens.

The benchmarks published with the release show 42.7 on HLE (out of 60.0 maximum), 87.9 on Terminal Bench 2.1, 62.7 on DeepSWE, and 83.3 on Cybergym. On SWE-bench Verified, the model ranked #2 out of 82 submissions at 96.40%. NVIDIA Build catalog now lists the DeepSeek-V4-Pro-0813 model card.

Google shipped Gemini 3.7 Flash on the same day. The changelog confirms August 13, 2026 as the GA date. Google positioned it for coding and agent workflows. FrontierCode 1.1 score is 43.6%, DeepSWE v1.1 is 65.3%, and WebDev Arena Elo is 1588.

The blog framed this as algorithmic improvements rather than a bigger model. No parameter count was disclosed.

Alibaba released Qwen3.8-Max as a 2.4T-parameter flagship with 95B active parameters. The weights are downloadable and runnable as open-weight software. Reuters reported on August 7 that Alibaba plans to charge large users, while CNBC noted the weights remain available for download and local execution.

ByteDance leadership admitted its LLMs lag overseas labs, but Doubao remained competitive. Doubao has 382M monthly active users, Qwen 167M, and DeepSeek approximately 130M.

OpenAI published first results for Jalapeño, emphasizing speed across frontier workloads. The comparison included GPT-OSS at 120B, DeepSeek R1 at 670B, and Kimi K2.5 at 1T.

Bloomberg published a chart tracking the US-China AI race, arguing the US lead is narrowing. The chart snippet includes Grok 4.5 at 81.6%.


Scoreboard

The table below shows confirmed claims with dates, numbers, and URLs where available.

SystemMetricValueSource
DeepSeek V4 ProContext / output1M / 384K tokensapi-docs.deepseek.com
DeepSeek V4 ProPricing$1.32 in / $3.96 out per MReuters Aug 13
DeepSeek V4 ProHLE42.7 / 60.0api-docs.deepseek.com
DeepSeek V4 ProTerminal Bench 2.187.9api-docs.deepseek.com
DeepSeek V4 ProDeepSWE62.7api-docs.deepseek.com
DeepSeek V4 ProCybergym83.3api-docs.deepseek.com
DeepSeek V4 ProSWE-bench Verified rank#2 of 82, 96.40%vals.ai
Gemini 3.7 FlashGA dateAugust 13, 2026ai.google.dev changelog
Gemini 3.7 FlashFrontierCode 1.143.6%deepmind.google model card
Gemini 3.7 FlashDeepSWE v1.165.3%deepmind.google model card
Gemini 3.7 FlashWebDev Arena Elo1588deepmind.google model card
Qwen3.8-MaxParameters (total / active)2.4T / 95BReuters Aug 7
DoubaoMonthly active users382MCaixin Aug 7
QwenMonthly active users167MCaixin Aug 7
DeepSeekMonthly active users~130MCaixin Aug 7
OpenAI JalapeñoGPT-OSS comparison120Bopenai.com
OpenAI JalapeñoDeepSeek R1 comparison670Bopenai.com
OpenAI JalapeñoKimi K2.5 comparison1Topenai.com
NVIDIA Nemotron 3 Nano Omni 30B A3BThroughput323 tokens/secgmicloud.ai blog Aug
homelab RTX 3060 12GBOllama throughput69.34 tok/s (eval_count 64, eval_s 0.923)local measurement

Open-weight landscape

Hugging Face published a summer 2026 state-of-open-models report. The analysis of Chinese releases shows 59% use Apache 2.0 and 22% use MIT licenses.

Rest of World reported that the Chinese open-weight model surge is reshaping Silicon Valley debate. CNBC covered the U.S. open-weight push as part of the China competition.

CNBC also reported that Alibaba and Meta continue the open-weight AI laptop model trend, citing the downloadable and runnable nature of Qwen3.8-Max.


Three practitioner lines

Act: Test DeepSeek V4 Pro on SWE-bench Verified tasks if your workload matches the agentic or terminal benchmark domains; the pricing is now flagship-tier, so compare cost per solved ticket against Gemini 3.7 Flash and your existing pipeline.

Watch: Open-weight 2T+ MoE models (Qwen3.8-Max at 2.4T total, 95B active) are now downloadable and runnable; monitor inference cost and latency on your hardware before committing to API-only workflows.

Ignore: Parameter count headlines without active-parameter or throughput numbers; the homelab RTX 3060 measurement (69.34 tok/s, eval_count 64) and the NVIDIA Nemotron report (323 tok/s) show that architectural choices and quantization matter more than headline numbers for local execution.


⚠️ 加工엔진(멀티모델 검증)이 主이지 원료(밖)나 저자(안)가 主 아님. ①출처실재·②순서강건 교차검증만 자동. 최종 게시·맥락판단은 사람.