Ledger on GitHub. Hub here. Branches elsewhere.
Deep cut
Deep cut
2026-09-03 07:02 · 밖→안(팩토리 加工)→밖 · 다소스intake+순서강건교차검증+링크게이트 · 게시 前 사람 승인
Lab notes: August 2026 frontier release wave
DeepSeek shipped V4 Pro on August 13, 2026. The official release supersedes the preview and adds stronger agentic capability. The model is available via API, web interface, and mobile app. Context window is 1M tokens with 384K maximum output.
Pricing moved to flagship levels: $1.32 per million input tokens and $3.96 per million output tokens.
The benchmarks published with the release show 42.7 on HLE (out of 60.0 maximum), 87.9 on Terminal Bench 2.1, 62.7 on DeepSWE, and 83.3 on Cybergym. On SWE-bench Verified, the model ranked #2 out of 82 submissions at 96.40%. NVIDIA Build catalog now lists the DeepSeek-V4-Pro-0813 model card.
Google shipped Gemini 3.7 Flash on the same day. The changelog confirms August 13, 2026 as the GA date. Google positioned it for coding and agent workflows. FrontierCode 1.1 score is 43.6%, DeepSWE v1.1 is 65.3%, and WebDev Arena Elo is 1588.
The blog framed this as algorithmic improvements rather than a bigger model. No parameter count was disclosed.
Alibaba released Qwen3.8-Max as a 2.4T-parameter flagship with 95B active parameters. The weights are downloadable and runnable as open-weight software. Reuters reported on August 7 that Alibaba plans to charge large users, while CNBC noted the weights remain available for download and local execution.
ByteDance leadership admitted its LLMs lag overseas labs, but Doubao remained competitive. Doubao has 382M monthly active users, Qwen 167M, and DeepSeek approximately 130M.
OpenAI published first results for Jalapeño, emphasizing speed across frontier workloads. The comparison included GPT-OSS at 120B, DeepSeek R1 at 670B, and Kimi K2.5 at 1T.
Bloomberg published a chart tracking the US-China AI race, arguing the US lead is narrowing. The chart snippet includes Grok 4.5 at 81.6%.
Scoreboard
The table below shows confirmed claims with dates, numbers, and URLs where available.
| System | Metric | Value | Source |
|---|---|---|---|
| DeepSeek V4 Pro | Context / output | 1M / 384K tokens | api-docs.deepseek.com |
| DeepSeek V4 Pro | Pricing | $1.32 in / $3.96 out per M | Reuters Aug 13 |
| DeepSeek V4 Pro | HLE | 42.7 / 60.0 | api-docs.deepseek.com |
| DeepSeek V4 Pro | Terminal Bench 2.1 | 87.9 | api-docs.deepseek.com |
| DeepSeek V4 Pro | DeepSWE | 62.7 | api-docs.deepseek.com |
| DeepSeek V4 Pro | Cybergym | 83.3 | api-docs.deepseek.com |
| DeepSeek V4 Pro | SWE-bench Verified rank | #2 of 82, 96.40% | vals.ai |
| Gemini 3.7 Flash | GA date | August 13, 2026 | ai.google.dev changelog |
| Gemini 3.7 Flash | FrontierCode 1.1 | 43.6% | deepmind.google model card |
| Gemini 3.7 Flash | DeepSWE v1.1 | 65.3% | deepmind.google model card |
| Gemini 3.7 Flash | WebDev Arena Elo | 1588 | deepmind.google model card |
| Qwen3.8-Max | Parameters (total / active) | 2.4T / 95B | Reuters Aug 7 |
| Doubao | Monthly active users | 382M | Caixin Aug 7 |
| Qwen | Monthly active users | 167M | Caixin Aug 7 |
| DeepSeek | Monthly active users | ~130M | Caixin Aug 7 |
| OpenAI Jalapeño | GPT-OSS comparison | 120B | openai.com |
| OpenAI Jalapeño | DeepSeek R1 comparison | 670B | openai.com |
| OpenAI Jalapeño | Kimi K2.5 comparison | 1T | openai.com |
| NVIDIA Nemotron 3 Nano Omni 30B A3B | Throughput | 323 tokens/sec | gmicloud.ai blog Aug |
| homelab RTX 3060 12GB | Ollama throughput | 69.34 tok/s (eval_count 64, eval_s 0.923) | local measurement |
Open-weight landscape
Hugging Face published a summer 2026 state-of-open-models report. The analysis of Chinese releases shows 59% use Apache 2.0 and 22% use MIT licenses.
Rest of World reported that the Chinese open-weight model surge is reshaping Silicon Valley debate. CNBC covered the U.S. open-weight push as part of the China competition.
CNBC also reported that Alibaba and Meta continue the open-weight AI laptop model trend, citing the downloadable and runnable nature of Qwen3.8-Max.
Three practitioner lines
Act: Test DeepSeek V4 Pro on SWE-bench Verified tasks if your workload matches the agentic or terminal benchmark domains; the pricing is now flagship-tier, so compare cost per solved ticket against Gemini 3.7 Flash and your existing pipeline.
Watch: Open-weight 2T+ MoE models (Qwen3.8-Max at 2.4T total, 95B active) are now downloadable and runnable; monitor inference cost and latency on your hardware before committing to API-only workflows.
Ignore: Parameter count headlines without active-parameter or throughput numbers; the homelab RTX 3060 measurement (69.34 tok/s, eval_count 64) and the NVIDIA Nemotron report (323 tok/s) show that architectural choices and quantization matter more than headline numbers for local execution.
⚠️ 加工엔진(멀티모델 검증)이 主이지 원료(밖)나 저자(안)가 主 아님. ①출처실재·②순서강건 교차검증만 자동. 최종 게시·맥락판단은 사람.