DeepSeek-V4 Pro 正式版发布

智能体基准较预览版跃升近五倍,V4 系列全部转正

DeepSeek 将 V4 Pro 更新为 0813 正式版:1.6T 总参 MoE、1M 上下文、最大输出 38.4 万 token,首次原生支持图像推理;DeepSWE 从预览版 12.8 升至 62.7,超过 Claude Opus 4.8,部分智能体基准逼近甚至反超 Claude Fable 5。同日发布开源 Agent 开发框架 DeepSeek Harness v0.1。

时间2026 年 8 月 13 日 级别A · 行业级 组织DeepSeek(深度求索) 状态已核验 · 3 个来源
编辑插图:深蓝背景上一条代码流水线从低处跃升到高处,节点依次点亮
AI Chronicle 原创插图:基准跃升的阶梯,对应 V4 Pro 正式版的智能体能力跨越。 AI Chronicle

2026 年 8 月 13 日,DeepSeek 把 V4 Pro 更新为 0813 正式版。距离 4 月 24 日预览版上线整整 111 天。这 111 天里,V4-Flash 在 7 月 31 日先转正,Pro 的正式版成了 V4 系列最后一块拼图——但真正让这次更新成为新闻的,不是「转正」本身,而是一组数字:DeepSWE 从预览版的 12.8 跳到 62.7。

12.8 是预览版最大的软肋。4 月发布时,V4 Pro 的推理与价格都站得住,唯独智能体能力明显落后于同期旗舰——DeepSWE 12.8 这个分数,放在任何对比表里都刺眼。0813 正式版把这条短板补成了卖点:62.7 超过 Claude Opus 4.8 的 58.0,Terminal Bench 2.1 从 72.1 升到 87.9,逼近 Fable 5 的 88.0;Cybergym 从 52.7 升到 83.3。需要带着厂商自述的边界读,但方向是清楚的:DeepSeek 用一次正式版完成了智能体能力的代际跃升。

规格上,0813 与预览版架构基本一致:1.6T 总参、每 token 激活约 49B,上下文 1M、最大输出 38.4 万 token。新增的两件事更值得注意:一是首次原生支持图像推理,预览版是纯文本模型;二是同时兼容 OpenAI 与 Anthropic 两套 API 生态,Responses API 与 Codex 等智能体工具可以直接接入。对开发者来说,这意味着「换模型不换代码」的迁移成本被进一步压低。

同一天发布的 DeepSeek Harness v0.1(MIT 许可)容易被当成边角新闻,但它可能比模型本身更说明问题。Harness 是面向智能体编排的工程化框架——把竞争从「模型层」推进到「Agent 开发层」。当模型能力趋近时,框架、工具链与生态会成为下一个战场。DeepSeek 选择开源 Harness,等于在说:模型你们可以随便用,工具链我们也一起给。

V4 系列至此全部转正。对开发者,这是一个 1M 上下文、支持图像推理、双 API 生态、价格仍具优势的旗舰选项;对行业,DeepSWE 62.7 成为新的参照点——中国模型第一次在智能体基准的头部区间站稳。官方预告的「API 大幅调价」还没落地,但以 DeepSeek 的定价历史,价格战大概率只是时间问题。

On August 13, 2026, DeepSeek promoted V4 Pro to the 0813 GA build—111 days after the preview launched on April 24. In those 111 days, V4-Flash went GA on July 31, making the Pro GA the last piece of the V4 line. But what made this update news was not the GA itself; it was one set of numbers: DeepSWE jumping from 12.8 in preview to 62.7.

12.8 was the preview's biggest weakness. When V4 Pro launched in April, its reasoning and price held up, but agent capability clearly lagged contemporary flagships—a DeepSWE of 12.8 stood out on any comparison table. The 0813 GA turned that weakness into a selling point: 62.7 beats Claude Opus 4.8's 58.0, Terminal Bench 2.1 rose from 72.1 to 87.9 (near Fable 5's 88.0), and Cybergym from 52.7 to 83.3. Read with the vendor-self-reported caveat, but the direction is clear: DeepSeek completed a generational agent leap in one GA release.

On specs, 0813 keeps the preview architecture: 1.6T total, ~49B active per token, 1M context, 384K max output. Two additions matter more: first native image reasoning (the preview was text-only), and compatibility with both OpenAI and Anthropic API ecosystems, so Responses API and agent tools like Codex plug in directly. For developers, the migration cost of "switch model, keep code" drops further.

The DeepSeek Harness v0.1 (MIT) released the same day is easy to file as a side note, but it may say more than the model itself. Harness is an engineering framework for agent orchestration—extending competition from the model layer to the agent development layer. When model capability converges, frameworks, toolchains, and ecosystems become the next battlefield. DeepSeek open-sourcing Harness is a way of saying: use the models freely, we give you the toolchain too.

The V4 line is now fully GA. For developers, this is a flagship option with 1M context, image reasoning, dual API ecosystems, and still-aggressive pricing; for the industry, DeepSWE 62.7 is a new reference point—a Chinese model standing firmly in the top band of agent benchmarks for the first time. The promised "big API price cuts" have not landed yet, but given DeepSeek's pricing history, the price war is probably a matter of time.

展开完整事件档案人物、主题、模型与产品
人物
模型
deepseek-v4-pro
产品
来源

原始资料

  1. 01DeepSeek API 文档更新(V4-Pro-0813)DeepSeek · official
  2. 02DeepSeek Harness 开源仓库GitHub · official
  3. 0321财经:DeepSeek V4 Pro 正式版发布报道21财经 · report

试试搜索