dropkit.contents

Ledger on GitHub. Hub here. Branches elsewhere.

← Drops · 2026-09-02

Deep cut

Deep cut

2026-09-02 21:01 · 밖→안(팩토리 加工)→밖 · 다소스intake+순서강건교차검증+링크게이트 · 게시 前 사람 승인

Publication briefing: GLM-5.3, DeepSeek V4 Pro, Muse Glimmer 30B, Jalapeño infra, homelab baseline

Two open-weight MoE models crossed 700B total parameters this cycle. One agentic benchmark crossed 80% on SWE-bench Verified. One Apache 2.0 open-weight model at 30B reached 76.0 on the same benchmark. One inference accelerator claims sub-latency and power gains on models from 120B to 1T. One homelab RTX 3060 measurement anchors the column.

GLM-5.3

GLM-5.3 is an open-weight MoE model for agentic work with 753B total params, 40B active params, and up to 1M-token context. Specs are independently echoed by a second model tracker.

Pricing and throughput were reported around $1.20 input / $4.00 output per 1M tokens, approximately 280 tok/s, TTFT 0.432s. These figures come from third-party reporting only.

DeepSeek V4 Pro

DeepSeek V4 Pro was reported as a 1.6T-parameter MoE with 49B active params and 1M context. Official-style pricing was reported at $0.435 input / $0.87 output per 1M tokens. An external guide echoes the same list price.

The model was tied to vendor-reported benchmark claims including SWE-bench Verified 80.6%, GPQA Diamond 90.1%, and LiveCodeBench 93.5%. These are explicitly vendor-reported. DeepSeek V4 Pro 0813 was listed at about 1.9 s average total latency in a third-party API speed snapshot.

Muse Glimmer 30B

Muse Glimmer 30B was reported as a 30B Apache 2.0 open-weight model with 76.0 SWE-Bench Verified. The benchmark result is echoed by a second source.

Jalapeño

Jalapeño reportedly delivered 1.5–1.9× more AI work per watt and 1.7–3.6× lower end-to-end latency on models from 120B to 1T. This is a systems and infrastructure claim only.

Table

Model / systemTypeTotal paramsActive paramsContextPricing ($/1M tok, in/out)ThroughputLatencyBenchmarkURL
GLM-5.3open-weight MoE753B40B1M$1.20 / $4.00~280 tok/sTTFT 0.432shttps://lmstudio.ai/models/glm-5.3 ; https://openllmstack.com/models/glm-5-3/ ; https://whatllm.org/models/glm-5-3 ; https://www.qubrid.com/blog/glm-53-is-here-full-benchmark-breakdown-architecture-pricing
DeepSeek V4 ProMoE1.6T49B1M$0.435 / $0.87~1.9 s (0813)SWE-bench Verified 80.6%, GPQA Diamond 90.1%, LiveCodeBench 93.5% (vendor)https://www.morphllm.com/deepseek-v4 ; https://deepseek.ai/deepseek-v4 ; https://codersera.com/blog/deepseek-v4-complete-guide-2026/ ; https://www.progressiverobot.com/2026/08/13/open-weight-ai-models-2026/ ; https://www.ailatency.com/reports/daily-v2/2026-08-19.html
Muse Glimmer 30Bopen-weight, Apache 2.030BSWE-bench Verified 76.0https://developer.puter.com/ai/meta/muse-glimmer-30b/ ; https://aiweekly.co/alerts/meta-superintelligence-lab-open-sources-muse-glimmer-30b
Jalapeñoinference accelerator1.7–3.6× lower (120B–1T)1.5–1.9× per watthttps://paragraph.com/@twiata/this-week-in-all-things-ai-week-35-2026
homelab RTX 3060 12GB OllamaLOCAL MEASUREMENT69.52 tok/s0.849 eval_seval_count 59none (local)

Practitioner lines

Act: Compare GLM-5.3 and DeepSeek V4 Pro pricing against your current hosted inference bill if you run agent loops or long-context tasks. Run Muse Glimmer 30B locally if you need Apache 2.0 licensing and a 30B footprint that fits consumer GPUs.

Watch: Track whether Jalapeño latency and power claims hold across model families outside the 120B–1T range, and whether independent measurements appear. Watch for open-weight releases that challenge the 80.6% SWE-bench Verified ceiling.

Ignore: Pricing claims without cross-checks, benchmarks without dataset versions, and throughput numbers without TTFT or eval-time breakdowns. Ignore any prose that does not include a model name, parameter count, or URL.


⚠️ 加工엔진(멀티모델 검증)이 主이지 원료(밖)나 저자(안)가 主 아님. ①출처실재·②순서강건 교차검증만 자동. 최종 게시·맥락판단은 사람.