DeepSeek-V4-Flash 正式版发布

预览版转正并全线上调 API 定价,低价策略迎来转折点

DeepSeek 发布 V4-Flash 正式版(V4-Flash-0731),取代预览版:思考/非思考双模式、1M 上下文、Agent 能力增强与推测解码;8 月 6 日公告整体上调 API 定价,市场解读为低价策略转折与 V4-Pro 正式版信号。

时间2026 年 7 月 31 日 级别A · 行业级 组织DeepSeek(深度求索) 状态已核验 · 3 个来源
编辑插图:深蓝背景上的橙色价格标签与上升曲线
AI Chronicle 原创插图:价格曲线与标签,对应 V4-Flash 转正与 API 涨价。 AI Chronicle

2026 年 7 月 31 日,DeepSeek 发布 V4-Flash 正式版,版本号 V4-Flash-0731。预览版转正,思考/非思考双模式、1M 上下文、Agent 能力增强、推测解码——这些是技术清单。但让整个开发者社区坐直身子的,是六天后的另一份公告:8 月 6 日,DeepSeek 宣布整体上调 API 定价,「预计涨幅较大」。

两件事放在一起读,才构成完整的信号。Flash 转正说明模型进入稳定期,接口行为、速率限制不再频繁变动;涨价则说明 DeepSeek 决定不再用补贴式低价换市场。过去两年,DeepSeek 的 API 价格一直是行业地板,R1 时代「白菜价」甚至成为全球话题。当推理需求随 Agent 工作负载暴涨,地板价与算力账单之间的裂缝,终于到了需要修补的时候。

当前平峰价是:输入缓存命中 0.02 元、未命中 1 元,输出 2 元每百万 token,高峰时段翻倍。这个结构本身就在说话——缓存命中与未命中相差 50 倍,DeepSeek 在鼓励开发者把重复前缀缓存起来;高峰翻倍,则是在用价格把负载推向平峰。定价不再只是「便宜」,而是一套引导行为的工程工具。

对开发者,这是一次集体重算。依赖 DeepSeek API 的 Agent 产品、对话应用、云厂商转售服务,都要把「预计涨幅较大」写进成本模型。对行业,这是低价竞赛的转折信号:当最激进的低价玩家开始涨价,市场对「以价换量」策略的可持续性需要重新定价。8 月 2 到 4 日多家云平台上线 Flash 正式版,说明生态仍在扩张——涨价没有吓退渠道,反而让正式版的分发更整齐。

DeepSeek 没有解释涨价的财务细节,外界也不该替它编造原因。可以确认的是节奏:预览转正与涨价放在同一时间窗口,旧模型名退役的倒计时刚刚结束,新价格体系就位。低价时代没有突然结束,它只是被一份「预计涨幅较大」的公告,正式翻到了下一章。

On July 31, 2026, DeepSeek shipped the V4-Flash general availability build, version V4-Flash-0731. Preview to GA, thinking and non-thinking modes, a 1M-token context, stronger agent capability, speculative decoding—that was the technical list. What made the whole developer community sit up was another notice six days later: on August 6, DeepSeek announced a broad API price increase, "expected to be significant."

The two events only form a complete signal when read together. GA means the model entered a stable phase—interface behavior and rate limits stop shifting; the price increase means DeepSeek decided to stop buying market share with subsidized prices. For two years, DeepSeek's API pricing had been the industry floor; in the R1 era, "bargain prices" became a global talking point. As agent workloads drove inference demand up, the gap between floor prices and compute bills finally needed patching.

Current off-peak pricing: ¥0.02 per million input tokens on cache hits, ¥1 on misses, ¥2 per million output, doubling at peak hours. The structure itself speaks—a 50x gap between cache hit and miss is DeepSeek telling developers to cache repeated prefixes; peak doubling is price steering load toward off-peak hours. Pricing is no longer just "cheap"; it is an engineering tool that shapes behavior.

For developers, this was a collective re-budget. Agent products, chat apps, and cloud resellers on the DeepSeek API all had to write "expected to be significant" into their cost models. For the industry, it is a turning signal in the low-price race: when the most aggressive discounter starts raising prices, the sustainability of volume-for-price strategies gets repriced. Multiple cloud platforms onboarded the Flash GA between August 2 and 4, so the ecosystem was still expanding—the increase did not scare off channels; it made GA distribution tidier.

DeepSeek did not explain the financial details of the increase, and outsiders should not invent reasons for it. What is confirmable is the cadence: preview-to-GA and the price increase landed in the same window, right after the legacy model-name retirement countdown ended. The low-price era did not end abruptly; it was formally turned to the next chapter by a notice that said "expected to be significant."

展开完整事件档案人物、主题、模型与产品
人物
模型
deepseek-v4-flash
产品
来源

原始资料

  1. 01界面新闻:DeepSeek V4-Flash 正式版发布界面新闻 · report
  2. 02上海证券报:DeepSeek 上调 API 定价上海证券报 · report
  3. 03澎湃新闻:DeepSeek 涨价公告解读澎湃新闻 · report

试试搜索