Ledger on GitHub. Hub here. Branches elsewhere.
Deep cut
Deep cut
2026-09-02 21:01 · 밖→안(팩토리 加工)→밖 · 다소스intake+순서강건교차검증+링크게이트 · 게시 前 사람 승인
Publication briefing: GLM-5.3, DeepSeek V4 Pro, Muse Glimmer 30B, Jalapeño infra, homelab baseline
Two open-weight MoE models crossed 700B total parameters this cycle. One agentic benchmark crossed 80% on SWE-bench Verified. One Apache 2.0 open-weight model at 30B reached 76.0 on the same benchmark. One inference accelerator claims sub-latency and power gains on models from 120B to 1T. One homelab RTX 3060 measurement anchors the column.
GLM-5.3
GLM-5.3 is an open-weight MoE model for agentic work with 753B total params, 40B active params, and up to 1M-token context. Specs are independently echoed by a second model tracker.
Pricing and throughput were reported around $1.20 input / $4.00 output per 1M tokens, approximately 280 tok/s, TTFT 0.432s. These figures come from third-party reporting only.
DeepSeek V4 Pro
DeepSeek V4 Pro was reported as a 1.6T-parameter MoE with 49B active params and 1M context. Official-style pricing was reported at $0.435 input / $0.87 output per 1M tokens. An external guide echoes the same list price.
The model was tied to vendor-reported benchmark claims including SWE-bench Verified 80.6%, GPQA Diamond 90.1%, and LiveCodeBench 93.5%. These are explicitly vendor-reported. DeepSeek V4 Pro 0813 was listed at about 1.9 s average total latency in a third-party API speed snapshot.
Muse Glimmer 30B
Muse Glimmer 30B was reported as a 30B Apache 2.0 open-weight model with 76.0 SWE-Bench Verified. The benchmark result is echoed by a second source.
Jalapeño
Jalapeño reportedly delivered 1.5–1.9× more AI work per watt and 1.7–3.6× lower end-to-end latency on models from 120B to 1T. This is a systems and infrastructure claim only.
Table
Practitioner lines
Act: Compare GLM-5.3 and DeepSeek V4 Pro pricing against your current hosted inference bill if you run agent loops or long-context tasks. Run Muse Glimmer 30B locally if you need Apache 2.0 licensing and a 30B footprint that fits consumer GPUs.
Watch: Track whether Jalapeño latency and power claims hold across model families outside the 120B–1T range, and whether independent measurements appear. Watch for open-weight releases that challenge the 80.6% SWE-bench Verified ceiling.
Ignore: Pricing claims without cross-checks, benchmarks without dataset versions, and throughput numbers without TTFT or eval-time breakdowns. Ignore any prose that does not include a model name, parameter count, or URL.
⚠️ 加工엔진(멀티모델 검증)이 主이지 원료(밖)나 저자(안)가 主 아님. ①출처실재·②순서강건 교차검증만 자동. 최종 게시·맥락판단은 사람.