dropkit.contents

Ledger on GitHub. Hub here. Branches elsewhere.

← Drops · 2026-09-07

8 Updates From the First Week of August 2026

8 Updates From the First Week of August 2026

The first week of August 2026 saw several updates across major AI model providers and platforms. There is a noticeable focus on coding and agentic workflows, with several models being positioned for these specific use cases. Pricing structures remain a point of differentiation, with various tiers and cache mechanisms in play.

Gemini’s 3.7 Flash model has reached General Availability (GA). Its primary target applications are complex coding tasks and agentic workflows. Pricing for this model is set at $0.75 per 1 million input tokens and $3.75 per 1 million output tokens. This introductory pricing is scheduled to run through December 31, 2026. The official documentation and a related blog post confirm these details.

Grok has released version 4.6, with pricing tiered by context window size. For context windows below 200K, the cost is $2 per 1 million input tokens and $6 per 1 million output tokens. For context windows exceeding 200K, the price increases to $4 per 1 million input tokens and $12 per 1 million output tokens. The model supports a maximum context of 500K tokens. Grok 4.6 also reports a low Time To First Token (TTFT) of 6.98 seconds, which is noted as the lowest among comparisons for the first answer token. This pricing and latency data is independently reported and confirmed across several sources. Cached input for Grok 4.6 is priced at $0.50 per 1 million tokens.

Qwen’s 3.8-Max model has been launched, characterized as a frontier model with strong performance in terminal and coding tasks. It registers 86.6 on the Terminal-Bench 2.1 benchmark and has 2.4 trillion parameters. Its pricing is $2 per 1 million input tokens and $6 per 1 million output tokens. The pricing for Qwen 3.8-Max is consistently reported across multiple non-vendor write-ups. Additionally, cache read operations are $0.25, cache creation is $2.50, and a different cache read operation is $0.17.

GLM 5.3 is positioned as an open-weights model, focusing on coding and cyber-security applications. It has been associated with addressing 2,436 vulnerabilities and interacting with 269 open-source projects. Its API pricing is $1.40 per 1 million input tokens, $0.26 per 1 million cached input tokens, and $4.40 per 1 million output tokens. This pricing information is available in multiple summaries.

A broader trend observed in summer 2026 is the significant momentum in Chinese open models. A review indicates 178 Chinese model releases exceeding 20 billion parameters, with 59% licensed under Apache 2.0 and 22% under MIT, highlighting a permissive licensing approach.

While sensitive, some industry comparisons note the appearance of Claude Sonnet 5 in August 2026 API pricing comparisons at $2/$10 per 1 million tokens. ChatGPT’s flagship model is also reported to have transitioned to GPT-5.6 Luna in August 2026 comparisons. Qwen 3.8-Flash was described as an open-weight, multimodal model and a preview of Qwen 4, with 125 billion parameters plus 51 billion N-gram. AWS Bedrock AgentCore added web search for agents with citations and reached General Availability. OpenAI also released Codex Harness as an open-source execution framework for coding agents.

claimlabelorder (🟢 robust / ⚠️ sensitive)numbersURL
Gemini 3.7 Flash GA targets coding and agentic workflowsGemini🟢 robust$0.75 input / $3.75 output per 1M; intro pricing through 2026-12-31https://ai.google.dev/gemini-api/docs/changelog
Gemini 3.7 Flash is framed for complex coding and multi-step executionGemini🟢 robustsame pricing; coding, agentic workflowshttps://blog.google/innovation-and-ai/models-and-research/gemini-models/introducing-gemini-3-7-flash/
Claude Sonnet 5 appears in August 2026 API pricing comparisonsClaude⚠️ sensitive$2 / $10 per 1Mhttps://compareai.today/blog/ai-api-pricing-august-2026
ChatGPT flagship moved to GPT-5.6 Luna in August 2026 comparisonsChatGPT⚠️ sensitiven/ahttps://thursdai.news/releases/2026-08
Grok 4.6 pricing is tiered by context windowGrok🟢 robust$2/$6 per 1M below 200K; $4/$12 above 200K; 500K contexthttps://vercel.com/ai-gateway/models/grok-4.6
Grok 4.6 latency is reported as low TTFT in a model comparison dashboardGrok🟢 robust6.98s TTFT; lowest time to first answer tokenhttps://artificialanalysis.ai/models/releases/grok-4-6
Grok 4.6 pricing is repeated in independent launch coverageGrok🟢 robust$2 input / $6 output per 1M; cached input $0.50https://codersera.com/blog/grok-4-6-launch-guide-2026/
Qwen3.8-Max launched as a frontier model with terminal/coding strength and aggressive pricingQwen⚠️ sensitive2.4T; 86.6 Terminal-Bench 2.1; $2/$6 per 1M tokenshttps://www.marktechpost.com/2026/08/03/alibaba-qwen-releases-qwen3-8-max/
Qwen3.8-Max pricing was repeated in multiple non-vendor writeupsQwen🟢 robust$2/$6 per 1M tokens; $0.25 cache read; $2.50 cache creation; $0.17 cache readhttps://www.marktechpost.com/2026/08/03/alibaba-qwen-releases-qwen3-8-max/
Qwen3.8-Flash was described as open-weight, multimodal, and a Qwen4 previewQwen⚠️ sensitive125B + 51B N-gramhttps://www.requesty.ai/blog/open-weight-frontier-august-2026-glm-qwen-hy4
GLM 5.3 was positioned as a coding/cyber-focused open-weights modelGLM⚠️ sensitive2,436 vulnerabilities; 269 open-source projectshttps://www.ventureatlas.org/news/2026-08-20-zhipu-glm-5-3-coding-cyber
GLM 5.3 API pricing was published in multiple summariesGLM🟢 robust$1.40 input; $0.26 cached input; $4.40 outputhttps://ofox.ai/blog/glm-5-3-benchmarks-access-2026/
AWS Bedrock AgentCore added web search for agents with citationsagents⚠️ sensitiveGAhttps://agentry.news/authors/1f7e1d07-84fc-4609-976e-0d875ed4ddb1
OpenAI released Codex Harness as an open-source execution framework for coding agentscoding⚠️ sensitiven/ahttps://agentry.news/authors/1f7e1d07-84fc-4609-976e-0d875ed4ddb1
Chinese open-model momentum in summer 2026 was broad and permissive-license heavysearch🟢 robust178 Chinese releases >20B; 59% Apache 2.0; 22% MIThttps://huggingface.co/blog/state-of-open-models-summer-2026

Act: Evaluate Grok 4.6 for applications sensitive to Time To First Token (TTFT) due to its reported low latency, and consider its tiered pricing model for cost-effective context window management. Watch: Monitor the consistent positioning of Gemini 3.7 Flash for coding and agentic workflows, as this indicates a clear strategic focus for Google in these domains. Ignore: Speculative API pricing comparisons for unreleased models like Claude Sonnet 5, given their sensitive nature and potential for change.