Ledger on GitHub. Hub here. Branches elsewhere.
LLM Pricing and Performance Briefing
LLM Pricing and Performance Briefing
This briefing summarizes recent information regarding large language model (LLM) pricing, performance, and strategic positioning. The focus is on API pricing and reported capabilities for coding and agentic workloads.
Gemini 3.7 Flash carries introductory API pricing. This model is positioned for coding and agent use cases. The stated rates are $0.75 input and $3.75 output per 1 million tokens. Google also lists Gemini 3.7 Flash in its agent platform pricing at these same introductory rates.
Claude Opus 4.6 maintains premium frontier pricing. This model is also positioned for coding and agent tasks. Its pricing is $5 input and $25 output per 1 million tokens. These Opus pricing figures are consistent across Anthropic’s news and product pages.
Grok 4.6 is priced as a more economical frontier option. A fast variant of Grok 4.6 is also available. Base pricing for Grok 4.6 is $2 input and $6 output per 1 million tokens. OpenRouter’s Grok 4.6 page confirms this base pricing and adds a cost for web-search calls: $5 per 1,000 calls.
The following table presents a public ledger of claims, labels, order, numbers, and URLs related to LLM developments.
| claim | label | order | numbers | URL |
|---|---|---|---|---|
| Gemini 3.7 Flash has introductory API pricing and is positioned for coding and agent use. | 🟢 robust | 1 | $0.75 input / $3.75 output per 1M tokens | https://ai.google.dev/gemini-api/docs/pricing |
| Google also lists Gemini 3.7 Flash in its agent platform pricing with the same introductory rates. | 🟢 robust | 2 | $0.75 / $3.75 per 1M tokens | https://cloud.google.com/gemini-enterprise-agent-platform/generative-ai/pricing |
| Claude Opus 4.6 keeps premium frontier pricing for coding and agents. | 🟢 robust | 3 | $5 input / $25 output per 1M tokens | https://www.anthropic.com/news/claude-opus-4.6 |
| Claude Opus pricing is also stated on Anthropic’s product page. | 🟢 robust | 4 | $5 / $25 per 1M tokens | https://www.anthropic.com/claude/opus |
| Grok 4.6 is priced as a cheaper frontier option with a fast variant available. | 🟢 robust | 5 | $2 input / $6 output per 1M tokens | https://x.ai/news/grok-4.6 |
| OpenRouter’s Grok 4.6 page reflects the same base pricing and adds web-search call pricing. | 🟢 robust | 6 | $2 / $6 per 1M tokens; $5 per 1K web-search calls | https://openrouter.ai/x-ai/grok-4.6 |
| Gemini 2.5 Flash-Lite was benchmarked as a strong latency/value option. | ⚠️ sensitive | 7 | 0.29–0.35s TTFT; 213.5 tokens/s; $0.10 / $0.40 per 1M | https://www.kunalganglani.com/blog/llm-latency-benchmark-optimization |
| Gemini 3.6 Flash appeared in benchmark coverage as a lower-cost, high-throughput option. | ⚠️ sensitive | 8 | $1.50 input / $7.50 output per 1M tokens | https://tech-insider.org/gemini-3-6-flash-launch-2026/ |
| Grok 4.6 was also covered in benchmark and pricing roundups as a strong agentic model. | ⚠️ sensitive | 9 | DFO 86.6; 500K context; $2 / $6 per 1M | https://designforonline.com/the-best-ai-models-so-far-in-2026/ |
| Claude, ChatGPT, and Gemini were compared on coding benchmarks, with Claude slightly ahead in the cited roundup. | ⚠️ sensitive | 10 | Claude Opus 4.8 88.6%; GPT-5.5 88.7%; Gemini 3.1 Pro 80.6% | https://tech-insider.org/claude-vs-chatgpt-vs-gemini-2026/ |
| OpenAI previewed an Ultrafast GPT-5.6 Sol tier claiming up to 14× speed on Cerebras. | ⚠️ sensitive | 11 | up to 14× | https://buttondown.com/JacopoCastellano/archive/ai-digest-2026-08-15/ |
| Cerebras reported GPT-5.6 Sol Ultrafast handling 2,500 questions in just over 11 hours. | ⚠️ sensitive | 12 | 2,500 questions; 11h+ | https://investors.cerebras.ai/news-releases/news-release-details/cerebras-powers-ultrafast-mode-openais-gpt-56-sol/ |
Act: Investigate the API pricing of Gemini 3.7 Flash for coding and agent workflows, given its stated introductory rates.
Watch: Monitor the continued development and pricing of Grok 4.6, particularly its agentic performance and web-search integration.
Ignore: Claims related to specific benchmark percentages or speed multiples without consistent, verifiable external validation.