Teams that route production traffic through Chinese APIs just lost a simple rule of thumb. In five days, DeepSeek raised some API tiers by as much as 1,100%, Alibaba open-weighted a 2.4-trillion-parameter flagship it had never released, and Zhipu shipped GLM-5.3 with a roughly 6x coding-bench jump on the same base model. The shared signal is not "China got cheaper." It is a shift from competing on floor price to competing on pricing power.
This piece stays inside the source brief: timeline, the RMB price sheet, Qwen specs, GLM benches, three strategies, a head-to-head table, disputed claims, and a six-step check. You should leave knowing which tier the 1,100% headline refers to, who can use Qwen3.8-Max commercially, and whether GLM-5.3 is a new foundation model.
01 DeepSeek price hike timeline: what landed, and when
The pain is operational. Headlines quote 11x, 1,100%, and 350% as if they were one number. A license rumor claimed US, EU, UK, and Korean users were banned from downloading Qwen weights. Both problems are solvable if you put dates and line items on one page.
| Date | Event |
|---|---|
| Jul 16, 2026 | Moonshot AI open-weights Kimi K3 (2.8T parameters), drawing US security scrutiny |
| Jul 30, 2026 | OpenAI cuts GPT-5.6 Luna by 80% |
| Aug 2–3, 2026 | Alibaba previews, then launches, Qwen3.8-Max as a hosted API |
| Aug 6–7, 2026 | OpenAI makes Luna the free default with unlimited text chats |
| Aug 10, 2026 | Meta releases Muse Glimmer (30B, Apache 2.0) and teases open weights for flagship Muse Spark 1.2 |
| Aug 12, 2026 | Alibaba publishes Qwen3.8-2.4T-A95B open weights on Hugging Face / ModelScope; xAI ships Grok 4.6 |
| Aug 13, 2026 | DeepSeek-V4-Pro goes GA and announces a price increase effective Aug 17; Google ships discounted Gemini 3.7 Flash |
| Aug 14, 2026 | Zhipu ships GLM-5.3, reusing GLM-5.2's 743B base |
| Aug 17, 2026, 00:00 Beijing | DeepSeek's new pricing takes effect |
Zoom out and the two sides of the same fight appear. Chinese labs raised prices and opened flagship weights. US labs cut prices and went free at the consumer layer in the same window. Chinese financial media has started calling the domestic cadence "three model updates a week," grouping DeepSeek, Alibaba, Zhipu, Moonshot's Kimi K3, and MiniMax H3 as a force that is "forcing a global repricing of the AI industry."
One lab prices time of day. One lab publishes weights with a revenue ceiling. One lab scales post-training only. Together they are defining who sets the price, not who undercuts it.
02 DeepSeek new rates, Qwen3.8 specs, and GLM-5.3 benches
Peak hours are 9am–12pm and 2pm–6pm Beijing time. Figures below are RMB per 1M tokens from DeepSeek's official announcement, cross-checked in the source brief against Wall Street CN, IT Home, and V2EX.
| Billing item (per 1M tokens) | Old | New off-peak | New peak | Peak increase |
|---|---|---|---|---|
| V4-Flash cache hit (input) | ¥0.02 | ¥0.05 | ¥0.10 | ~400% |
| V4-Flash cache miss (input) | ¥1.0 | ¥1.5 | ¥3.0 | 200% |
| V4-Flash output | ¥2.0 | ¥4.5 | ¥9.0 | 350% |
| V4-Pro cache hit (input) | ¥0.025 | ¥0.15 | ¥0.30 | ~1,100% |
| V4-Pro cache miss (input) | ¥3.0 | ¥4.5 | ¥9.0 | 200% |
| V4-Pro output | ¥6.0 | ¥13.5 | ¥27.0 | 350% |
The 1,100% headline applies to peak-hour cache-hit input — the tier that started closest to free. Output, which dominates most real bills, rose 350%. Independent cost modeling cited in the English brief found a realistic heavy-usage workload (roughly 84M tokens/month, mostly off-peak, half cache hits) sees a bill increase closer to 1.8x.
| Spec | Detail |
|---|---|
| Parameters | 2.4T total, 95B active per token (MoE, 512 experts, 10 routed + 1 shared) |
| Context window | 262,144 tokens native (open checkpoint), extendable to ~1.01M; hosted Max defaults to 1M |
| Release cadence | Preview Aug 2 → API live Aug 3 → open weights Aug 12 |
| API pricing (international) | $2/M input, $6/M output |
| License | Not Apache 2.0 — a custom "Qwen3.8-Max License" |
| Why it matters | First time Alibaba has open-weighted a Max-tier flagship; Qwen3.5 / 3.6 / 3.7 Max stayed API-only |
| Benchmark | GLM-5.2 | GLM-5.3 | Change |
|---|---|---|---|
| Terminal-Bench 3.0 | 4.6% | 28.3% | +23.7 pts |
| DeepSWE v1.1 | 46.2% | 66.9% | +20.7 pts |
| Agents' Last Exam (CLI) | 23.8% | 28.5% | +4.7 pts |
| CyberGym | 77.2% | 84.5% | +7.3 pts |
| AutomationBench | 26.2% | 48.2% | +22.0 pts |
These are Zhipu's own numbers. No independent third-party re-run has been published. GLM-5.3 still trails GPT-5.6 Sol (34.6%) and Claude Fable 5 (33.7%) on Terminal-Bench 3.0. It is a top open-weight result, not an outright frontier win.
Quote the cache-hit input line if you want the scare number. Quote output and cache-miss input if you want the bill.
03 DeepSeek time-of-day pricing, Alibaba's license, Zhipu's post-training bet
DeepSeek made capacity visible. The easy misread is "China's cheapest model finally caved to margin pressure." The structure reads more like a company putting compute scarcity on the price sheet. Flat, always-cheap pricing worked as customer acquisition while GPUs kept up. Once usage grew exponentially and capacity did not, "encouraging more flexible workload scheduling" became corporate-speak for "peak-hour compute is scarce, shift the load yourself." One detail international coverage mostly missed: at peak hours, DeepSeek's official API is now higher than several third-party resellers (GMI Cloud, Novita, and others currently list V4 Pro below the new official peak). The assumption that the official API is always the cheapest way to run DeepSeek has been broken.
Alibaba bought mindshare and kept the revenue ceiling. Publishing the 2.4T checkpoint is not a blank check. The custom license — not the Apache 2.0 used for smaller Qwen models — requires any "Model-as-a-Service" or "AI Work Assistant" business earning over $50 million in any 12-month period to negotiate a separate commercial license. Products with 100M+ monthly active users or $20M+ in monthly revenue must prominently display the model name. That is a different bet from Meta's Muse Glimmer, which ships under unrestricted Apache 2.0. Claims that the license bans downloads from the US, EU, UK, and South Korea are false. The published text has no geographic clause.
GLM-5.3 did not retrain the base. Same 743B-parameter foundation as GLM-5.2. No pretraining rerun. Terminal-Bench 3.0 moved from 4.6% to 28.3% — roughly 6x — by scaling reinforcement-learning environments in post-training. As pretraining scaling laws show diminishing returns, post-training RL is becoming an independent lever with a lower cost floor than a new foundation model. Mid-tier labs can close gaps on agentic and coding benches without OpenAI-scale pretraining budgets.
China labs · Aug 12–17 2026
DeepSeek TOD pricing · cache-hit input ~1100% peak
Qwen3.8 2.4T-A95B weights + custom license
GLM-5.3 same 743B base · post-training only
US labs Luna -80% then free default · Gemini 3.7 Flash cut
Signal pricing power, not floor price
- V4-Pro cache-hit input: ¥0.025 → ¥0.30 peak, about 1,100%.
- Qwen3.8-Max: 2.4T total / 95B active; international API $2 / $6; custom license.
- GLM-5.3 Terminal-Bench 3.0: 4.6% → 28.3% (vendor self-test); still behind Sol 34.6% and Fable 5 33.7%.
04 Is DeepSeek still the cheapest frontier model? Six checks
RMB-to-USD conversion in the brief is about ¥7.15/$1. Treat it as approximate.
| Model | Input (per 1M tokens) | Output (per 1M tokens) | Open weights? |
|---|---|---|---|
| DeepSeek V4-Pro (peak) | ¥9.0 (~$1.26) | ¥27.0 (~$3.78) | No |
| DeepSeek V4-Pro (off-peak) | ¥4.5 (~$0.63) | ¥13.5 (~$1.89) | No |
| Qwen3.8-Max (international API) | $2.00 | $6.00 | Yes (custom license) |
| OpenAI GPT-5.6 Luna | $0.20 | $1.20 | No |
| Claude Opus 5 (implied, per Alibaba's comparison ratio) | ~$5.00 | ~$25.00 | No |
Off-peak DeepSeek V4-Pro is still well below Claude Opus 5. It is no longer the outright cheapest option. Qwen3.8-Max international pricing and OpenAI's Luna both undercut DeepSeek's off-peak rate. "Chinese model = cheapest model" held for most of 2025 and early 2026. It is not a safe assumption now.
What the brief marks as disputed or unverified:
- The 1,100% headline is accurate and incomplete. It is peak cache-hit input only. Output rose 350%.
- Zhenwu M890 / Pangu AL128 claims appear in Chinese financial outlets as a domestic-silicon inference stack. They have not been confirmed by Alibaba technical docs or third-party benches. Treat as vendor-adjacent, unverified.
- GLM-5.3's reported "serious vulnerability" in Cursor comes from VentureBeat and Zhipu. Technical details are unpublished. Read it as vendor-sourced, not independently audited.
- Reports that China's Ministry of Commerce may prepare retaliatory export controls are speculative media reports, not an official announcement.
Six checks after you finish the headlines:
- Open DeepSeek's official sheet. Split cache-hit input, cache-miss input, and output. Peak is 9am–12pm and 2pm–6pm Beijing. The new rates start Aug 17, 00:00 Beijing.
- Map each headline to one line item. ~1,100% is V4-Pro peak cache-hit input. 350% is output. 200% is cache-miss input.
- Read the Qwen3.8-Max License file, not the announcement thread. No geographic ban. Separate commercial license above $50M trailing 12-month MaaS / AI Work Assistant revenue. Prominent model-name display above 100M MAU or $20M monthly revenue.
- Treat GLM benches as vendor self-tests. Same 743B base. Terminal-Bench 3.0 4.6% → 28.3%. Still behind Sol and Fable 5. No third-party re-run yet.
- Build your own matrix. DeepSeek off-peak / peak, Qwen $2 / $6, Luna $0.20 / $1.20, Claude Opus 5 ~$5 / $25. Do not default to "China is cheapest."
- Schedule the workload. Off-peak plus decent cache hits can land near 1.8x. Peak plus cache misses can hit the headline. If you self-host weights, price the custom license and the node, not just the download.
Primary sources (re-open after publication and confirm the live text):
05 DeepSeek pricing FAQ: bills, licenses, and where to run agents
FAQ
- Is DeepSeek still cheaper than GPT-5.6 or Claude after the hike? Off-peak is still cheaper than Claude Opus 5. It is no longer the single cheapest option. Luna at $0.20 / $1.20 and Qwen3.8-Max international at $2 / $6 undercut DeepSeek's new off-peak rates on at least one dimension. DeepSeek remains relatively cheap for a frontier-class model. It is not the outright cheapest.
- Can I use Alibaba's Qwen3.8-Max open weights for free in a commercial product? Yes for most cases. Personal projects and internal enterprise use are unaffected. The catch is a MaaS or AI Work Assistant business that earned over $50 million in any consecutive 12-month period. That tier needs a separate commercial license.
- Is Qwen3.8-Max banned or restricted for US, EU, or UK users? No. The published license has no geographic restriction. Limits are revenue-based, not location-based.
- What is actually different between GLM-5.3 and GLM-5.2? Nothing at the base-model level. Both use the same 743-billion-parameter foundation. The roughly 6x Terminal-Bench 3.0 gain comes from scaling reinforcement learning in post-training.
- Will Meta actually open-source its flagship, not just Muse Glimmer? Not yet. Muse Glimmer is a 30B distilled model. Mark Zuckerberg has said open weights for closed Muse Spark 1.2 are coming "soon." Treat that as a stated intention, not a fact on disk.
There is also a geopolitical layer the brief names carefully. Moonshot's Kimi K3 already drew US security scrutiny. Some analysts read Alibaba's 2.4T drop in this window as a move to lock in international mindshare and a "technological parity" narrative before any potential regulatory tightening. That is an informed interpretation, not a confirmed fact. English-language tech press has mostly covered these releases as isolated product news rather than as a coordinated national pattern.
If your stack only chases the official cheap API, peak TOD pricing will hit you. If you download 2.4T MoE weights without a stable node, you do not have a production environment. Shared GPU clouds also blur isolation, license display, and 7×24 logs. For teams that need a full macOS environment for Cursor, local open-weight trials, and iOS CI/CD — with nodes that stay online — CALMVPS bare-metal Mac Mini rental is usually the stronger fit: dedicated Apple Silicon, monthly elasticity, delivery in about 120 seconds. See the CALMVPS pricing page.
Sources: DeepSeek's official pricing announcement, cross-checked against Wall Street CN, IT Home, AIGC.cn, and V2EX; Alibaba's official Qwen repositories (Hugging Face / ModelScope) and South China Morning Post reporting on license terms; Zhipu (Z.ai) GLM-5.3 technical page, plus VentureBeat and StableLearn; Meta AI Research's official blog and VentureBeat on Muse Glimmer; Yicai and Sohu Finance on the pacing of China's open-weight cycle. Pricing, license terms, and benches reflect public information as of publication. Verify the latest official text before republishing. Domestic-chip claims, the Cursor vulnerability report, and export-control rumors remain unverified in the brief.