September 19, 2026
Chinese AI providers replaced discounts with tiers: Qwen3.7-flash pricing jumps 6.7-fold
Price per token no longer decides anything: the market compares the full cost of a task.

The same 657 office tasks cost $14 on DeepSeek-V4 Flash and about $1000 on Anthropic Opus 4.8.
ASPI's measurement from 21.07.2026 explains the gap. In agentic tasks, token consumption can run up to 3500 times higher than in a regular request, so fractions of a cent per token multiply into hundreds of dollars for a single task.
Vinnie Wu, head of Bank of America's Asian equity strategy, described the shift on 18.09.2026: the Chinese LLM price war is moving from blanket discounts to tiered rates. The market must now be measured by the cost of a completed task, not the price per token.
Tiers instead of discounts. In Alibaba Model Studio, the price tier is set by the total number of input tokens in a request, and the entire request is charged at that tier's rate. A 100K-token request above the 32K threshold moves entirely to the second tier: for Qwen3.7-flash, a million input tokens rises from $0.03 to $0.2, 6.7 times higher (Alibaba Cloud Model Studio docs, 18.09.2026).
The bill is cut even without changing models. DeepSeek keeps off-peak rates at exactly half the peak rate: peak hours run from 04:00 to 07:00 and from 09:00 to 13:00 MSK on weekdays, while all other hours, weekends and Chinese holidays are charged at half price (DeepSeek API docs, 19.09.2026). A million deepseek-flash input tokens costs $0.15 in that window instead of $0.30.
Discounts on top of the rate. Batch inference in Model Studio cuts input and output costs by 50%, while a context-cache hit costs 10% of the regular input rate. They do not stack; an explicit cache write costs 125% instead (Alibaba Cloud Model Studio docs, 18.09.2026).
Token prices kept falling nonetheless. The Silicon Data index showed $2.04 on May 31, 2026, $1.45 at the end of July and $1.16 on August 6–8, the year's low. Increases came in parallel: from February 12, 2026, Zhipu raised GLM-5 by at least 30% within China and by 67–100% for overseas API access, while Alibaba, ByteDance and two other major Chinese clouds moved their pricing to tiers.
Basic inference will keep getting cheaper, while BofA reserves a premium for models with reasoning, long context and agentic capabilities.
Source
