Ledger on GitHub. Hub here. Branches elsewhere.
drop-2026-09-04-kr-pm
5 Confirmed Cost-Efficiency Items from the Latest Lab Scan
Recent data points provide insights into the pricing and performance of prominent large language models (LLMs). This note synthesizes confirmed figures from various providers.
The Grok 4.6 model presents a confirmed pricing structure of $2 input per 1M tokens and $6 output per 1M tokens. Cache reads for Grok 4.6 are priced at $0.50 per 1M. Web search calls for Grok 4.6 are set at $5 per 1K calls. These figures offer a baseline for API integration and operational costs.
Claude Opus 5 documentation lists context, maximum output, price, and latency fields. This confirms the availability of detailed metadata for Opus 5. The pricing for Claude Opus 5 is confirmed at $5 input per 1M tokens and $25 output per 1M tokens. Cache reads for Claude Opus 5 are $0.50. Write operations for cache are $6.25 for 5 minutes and $10 for 1 hour. The latency for Claude Opus 5 is categorized as moderate.
GPT-5.6 Sol has a confirmed pricing of $5 input per 1M tokens and $30 output per 1M tokens.
Gemini models show confirmed pricing for different tiers. Gemini 3.8 Flash, 3.7 Flash, and 3.6 Flash are priced at $0.75 input per 1M and $3.75 output per 1M, with this rate valid through December 31, 2026. The Gemini Deep Research Agent is priced at $2 input per 1M and $12 output per 1M. Gemini Google Search grounding includes 5,000 free search requests per month, after which it costs $14 per 1,000 requests.
The following table summarizes the key confirmed data:
| claim | label | order (🟢 robust / ⚠️ sensitive) | numbers | URL |
|---|---|---|---|---|
| Grok 4.6 pricing and API rates | 🟢 confirmed | 🟢 robust | $2 input / $6 output per 1M; cache read $0.50 per 1M; web search $5 per 1K calls | https://x.ai/news/grok-4-6 |
| Claude Opus 5 docs with model metadata | 🟢 confirmed | 🟢 robust | context, max output, price, latency fields listed | https://platform.claude.com/docs/en/models/opus-5/overview |
| Gemini 3.8 Flash / 3.7 Flash / 3.6 Flash pricing | 🟢 confirmed | 🟢 robust | $0.75 / $3.75 per 1M input/output through Dec 31, 2026 | https://cloud.google.com/gemini-enterprise-agent-platform/generative-ai/pricing |
| Gemini Deep Research Agent pricing | 🟢 confirmed | 🟢 robust | Input $2 per 1M; output $12 per 1M | https://cloud.google.com/gemini-enterprise-agent-platform/generative-ai/pricing |
| Gemini Google Search grounding pricing | 🟢 confirmed | 🟢 robust | 5,000 free search requests/month, then $14 per 1,000 requests | https://ai.google.dev/gemini-api/docs/pricing |
| Claude Opus 5 pricing and cache rates | 🟢 confirmed | 🟢 robust | $5 input / $25 output per 1M; cache read $0.50; 5m write $6.25; 1h write $10 | https://platform.claude.com/docs/en/models/opus-5/overview |
| Claude Opus 5 latency band | 🟢 confirmed | 🟢 robust | Moderate latency | https://platform.claude.com/docs/en/models/opus-5/overview |
| GPT-5.6 Sol pricing | 🟢 confirmed | 🟢 robust | $5 input / $30 output per 1M | https://openai.com/index/gpt-5-6/ |
Practitioner Lines:
- Act: Leverage confirmed pricing for Grok 4.6 for API usage and Claude Opus 5 for detailed metadata in system design. Incorporate Gemini’s tiered pricing for Flash models and the Deep Research Agent into cost models.
- Watch: Monitor the long-term consistency of current pricing structures, particularly the Gemini Flash pricing which has an end date. Observe actual performance against the “moderate latency” claim for Claude Opus 5 in production environments.
- Ignore: Do not base long-term procurement decisions solely on partial or sensitive data points without independent verification.