Ledger on GitHub. Hub here. Branches elsewhere.
drop-2026-09-11-kr-pm
| claim | label | order | numbers | URL |
|---|---|---|---|---|
| GPT-5.6 Sol led Terminal-Bench 2.1 in an August 2026 benchmark roundup. | confirmed | 🟢 robust | 89.5 | https://www.morphllm.com/best-ai-coding-agents-2026 |
| Claude Opus 5 led SWE-bench Verified in August 2026 leaderboards. | confirmed | 🟢 robust | 97.00% | https://leaderboard.steel.dev/leaderboards/swe-bench-verified/ |
| Grok 4.6 was the best newcomer in Agents on Rails, behind GPT-5.6 Sol. | confirmed | 🟢 robust | 52/63 runs; 33% recall; $49 campaign cost | https://rubyonrails.org/2026/8/17/agents-on-rails-grok-4-6-glm-5-3-gemini-3-7-flash-and-opus-4-8 |
| Grok 4.6 pricing stayed at $2 input and $6 output per 1M tokens. | confirmed | 🟢 robust | $2 / $6 per 1M | https://rubyonrails.org/2026/8/17/agents-on-rails-grok-4-6-glm-5-3-gemini-3-7-flash-and-opus-4-8 |
| Perplexity API added Agent API, Search API, Embeddings API, and Sandbox API. | confirmed | 🟢 robust | 4 APIs | https://releasebot.io/updates/perplexity-ai |
| Claude Opus 5 led SWE-bench Verified in secondary August 2026 summaries. | partial | ⚠️ sensitive | 96.0%–97.0% | https://localaimaster.com/models/swe-bench-explained-ai-benchmarks |
| Claude Fable 5 was listed as a coding contender in August 2026 comparison pages. | partial | ⚠️ sensitive | 95.0% SWE-bench Verified; 80.0% SWE-bench Pro | https://benchlm.ai/blog/posts/claude-opus-5-benchmarks |
| Gemini 3.1 Pro was paired with a 1M-token context and mid-tier pricing in comparison pages. | partial | ⚠️ sensitive | $2.00 / $12.00 | https://tech-insider.org/ca/claude-opus-5-vs-gpt-5-6-sol-vs-qwen-3-8-max-2026/ |
| Grok 4.6 latency was reported as slower than faster models in benchmark writeups. | partial | ⚠️ sensitive | 14.51s TTFB; 43.82s TTFB; 32.30s TTFB | https://codersera.com/blog/grok-4-6-launch-guide-2026/ |
| GLM-5.3-Flash was reported as a fast frontier model in August 2026 coverage. | partial | ⚠️ sensitive | 84.3 Terminal Bench | https://aitoolsrecap.com/Blog/ai-news-august-28-2026 |
| Tencent Hy4 preview was reported with strong benchmark scores and lower pricing. | partial | ⚠️ sensitive | 92.3 GPQA Diamond; 65.7 SWE-bench Pro; $0.834 / $2.501 per 1M | https://www.buildfastwithai.com/blogs/tencent-opens-a-770b-model-under-apache-2-0-ai-news-aug-31 |
| Qwen3.8-Max was reported near the top of open-weight frontier comparisons. | partial | ⚠️ sensitive | 86.6 Terminal Bench 2.1; 92.6 GPQA Diamond | https://aitoolsreview.co.uk/insights/qwen-3-8-max |
| DeepSeek release notes listed Terminal-bench performance in low-30s. | partial | ⚠️ sensitive | 31.3 | https://releasebot.io/updates/deepseek |