dropkit.contents

Ledger on GitHub. Hub here. Branches elsewhere.

← Drops · 2026-09-05

drop-2026-09-05-kr-pm

AI Model API Pricing and Performance Snapshot: 5 Key Observations

This note outlines recent observations regarding AI model API pricing, performance, and feature-specific costs. Data is drawn from public API documentation and reported figures.

Claude Sonnet 5 API pricing has been noted. The initial introductory rate was later made permanent. Separate pricing for cache and web search is also documented.

Grok 4.5 API pricing was reported to include rates for input and output tokens. Cached-input pricing was also listed.

Tencent Hy4 preview performance metrics include reported throughput and P50 latency. Pricing for the Tencent Hy4 preview was also reported for input and output tokens.

Other models have also seen pricing reports. Gemini 3.1 Pro pricing has been reported for input and output tokens. Similarly, ChatGPT/GPT-5.5 Pro pricing was reported for its input and output tokens.

Qwen3.8-Max pricing included reported input and output token costs, with mention of cache discounts. DeepSeek V4 Pro pricing was reported for cache-miss input tokens and output tokens. Claude Code / Agent usage was described as separate from the shared subscription pool and is billed through API usage.

Here is a summary of the collected information:

claimlabelorder (🟢 robust / ⚠️ sensitive)numbersURL
Claude Sonnet 5 API pricing stayed at $2/$10 per 1M tokens, and Anthropic later said the introductory rate became permanent.🟢 robust🟢 robust$2 / $10https://www.anthropic.com/news/claude-sonnet-5
Claude Sonnet 5 docs list separate cache and web-search pricing alongside the token rates.🟢 robust🟢 robustcache read $0.20/M; cache write $2.50/M; 1h cache write $4/M; web search $10/1Khttps://platform.claude.com/docs/en/models/sonnet-5/overview
Grok 4.5 API pricing was reported at $2/$6 per 1M tokens with cached-input pricing also listed.🟢 robust🟢 robust$2 / $6; cached input $0.50/Mhttps://www.eesel.ai/blog/grok-4-5-pricing
Tencent Hy4 preview was reported at 26 tokens/s throughput with P50 latency of 3.77 seconds.🟢 robust🟢 robust26 tok/s; 3.77 s P50https://www.progressiverobot.com/2026/08/28/hy4-preview-tencent-open-weight-moe-1m-context/
Tencent Hy4 preview pricing was reported at roughly $0.834/M input and $2.501/M output tokens.🟡 sensitive⚠️ sensitive$0.834/M; $2.501/Mhttps://www.progressiverobot.com/2026/08/28/hy4-preview-tencent-open-weight-moe-1m-context/
Claude Code / Agent usage was described as separate from the shared subscription pool and billed through API usage.🟡 sensitive⚠️ sensitiveno numbers givenhttps://tech-insider.org/chatgpt-vs-gemini-vs-claude-pro-2026/
Gemini 3.1 Pro pricing was reported at $2/$12 per 1M tokens.🟡 sensitive⚠️ sensitive$2 / $12https://compareai.today/blog/ai-api-pricing-august-2026
ChatGPT/GPT-5.5 Pro pricing was reported at $30/$180 per 1M tokens.🟡 sensitive⚠️ sensitive$30 / $180https://compareai.today/blog/ai-api-pricing-august-2026
Qwen3.8-Max was reported at $2/$6 with cache discounts down to $0.25 per 1M implicit cache reads.🟡 sensitive⚠️ sensitive$2 / $6; $0.25https://www.marktechpost.com/2026/08/03/alibaba-qwen-releases-qwen3.8-max/
DeepSeek V4 Pro pricing was reported at $0.435 per 1M cache-miss input tokens and $0.87 per 1M output tokens.🟡 sensitive⚠️ sensitive$0.435; $0.87https://www.digitalapplied.com/blog/china-open-frontier-august-2026-scoreboard

Act: Evaluate the permanence of Claude Sonnet 5 API pricing, as Anthropic stated the introductory rate became permanent. This impacts long-term cost modeling.

Watch: Monitor the reported cache and web-search pricing for Claude Sonnet 5. These additional cost factors can significantly alter total API expenses.

Ignore: Do not base immediate operational decisions solely on the Tencent Hy4 preview pricing due to its ‘sensitive’ classification, indicating potential for change or verification needed.