DeepSeek-V4 预览版发布
开放权重模型迈入 1M 上下文与 Agent 编程
DeepSeek 开源 V4 Pro 与 Flash 预览版,把 1M 上下文、思考模式和 Agent 编程能力放进同一家族,并继续强调稀疏架构的成本效率。
2026 年 4 月 24 日,DeepSeek 在 API 文档中宣布 V4 Preview 上线并开放权重。页面一边强调“preview”,一边给出极其具体的迁移指令:保留原来的 base_url,把模型名改成 deepseek-v4-pro 或 deepseek-v4-flash。同一份说明还写明,旧的 deepseek-chat 与 deepseek-reasoner 将在 7 月 24 日退役。新一代尚未被称为正式版,旧入口的倒计时却已经开始。
Pro 与 Flash 把同一家族拆成两种取舍。官方资料列出的 Pro 为 1.6 万亿总参数、490 亿激活参数,面向更困难的推理与 Agent 编程;Flash 为 2840 亿总参数、130 亿激活参数,强调更快响应和更低价格。混合专家架构让总容量与每一步实际参与计算的参数分开,也让产品分层不再只是“同一个模型调低速度”。部署者选择的是两套不同的延迟、成本与能力边界。
所有官方服务把 100 万 token 上下文作为标准,并同时提供思考与非思考模式。这个组合针对的不是一次问答,而是更长的工作流:让 Agent 阅读仓库、调用工具、保留中间记录,再继续处理后续任务。窗口变长只解决“信息是否还能留在输入里”的问题;工具调用是否正确、较早约束是否被遵守、错误能否在多轮后被发现,仍需要实际任务验证。百万上下文是容器,不是长期工作的质量保证书。
发布页称 V4-Pro 在 Agent 编程基准中达到开放模型领先水平,并已接入 Claude Code、OpenCode 等工具。这些都是厂商陈述,评测脚本、提示模板和集成版本决定数字如何解释。预览阶段尤其如此:速率限制、接口行为和推荐模板都可能变化。开放权重让外部更早下载与复现,也让外部更早遇到尚未冻结的实现细节。
这次发布最真实的张力,藏在产品运维而非排行榜里。开发团队不能只问“新模型强不强”,还要决定何时修改模型名、怎样回归测试思考模式、Pro 与 Flash 如何分流,以及旧接口退役前能否完成迁移。preview 通常表示可以等待,退役通知却给等待本身设了截止日期。
DeepSeek-V4 没有宣告长期 Agent 已经被解决。它把百万上下文、双模式和两档稀疏模型放进一个可下载、可调用但仍在变化的版本,并用旧模型的退役日期推动真实用户进入。读懂这次发布,不只要看参数表,也要看那行迁移说明:前沿模型的迭代已经快到“预览”与“必须准备上线”可以同时成立。
On 24 April 2026, DeepSeek announced V4 Preview in its API documentation and released open weights. The page emphasized preview while giving a concrete migration instruction: keep the existing base_url, but change the model name to deepseek-v4-pro or deepseek-v4-flash. The same notice said that deepseek-chat and deepseek-reasoner would be retired on 24 July. The new generation had not been called generally available; the countdown for the old entry points had already begun.
Pro and Flash divided the family into two distinct tradeoffs. DeepSeek listed Pro at 1.6 trillion total parameters and 49 billion active, targeting harder reasoning and agentic coding. Flash had 284 billion total and 13 billion active, emphasizing faster responses and lower cost. Mixture-of-experts architecture separated stored capacity from the parameters used on each step. The product tiers were therefore not merely the same model at different speed settings. They carried different latency, cost, and capability boundaries.
Official services made a one-million-token context window standard and offered both thinking and non-thinking modes. The combination addressed work longer than a single answer: an agent reading a repository, calling tools, retaining intermediate records, and continuing across later tasks. A longer window solves only whether information can remain present in the input. Correct tool use, obedience to earlier constraints, and recovery from an error many steps later still require evaluation on real workflows. A million-token container is not a quality guarantee for long-horizon work.
The launch page described V4-Pro as leading open models on agentic-coding benchmarks and said the family integrated with tools including Claude Code and OpenCode. These were publisher claims. Evaluation scripts, prompt templates, and integration versions determine how the numbers should be read. That boundary matters even more in preview, when rate limits, interface behavior, and recommended templates may change. Open weights let outsiders download and reproduce earlier; they also expose outsiders earlier to implementation details that have not yet settled.
The release’s most concrete tension appeared in operations rather than on a leaderboard. A development team could not ask only whether the new model was stronger. It had to choose when to change names, how to regression-test thinking mode, how to route work between Pro and Flash, and whether migration could finish before the old endpoints disappeared. Preview usually suggests that waiting is reasonable. A retirement notice gives waiting a deadline.
DeepSeek-V4 did not announce that long-running agents had been solved. It placed million-token context, two modes, and two sparse model tiers into a downloadable and callable release that was still changing, then used the retirement date of older models to pull real users toward it. Understanding the event requires the migration line beside the parameter table. Frontier-model iteration had accelerated until “preview” and “prepare for production migration” could be true at the same time.
展开完整事件档案人物、主题、模型与产品
- 人物
- —
- 模型
- deepseek-v4-prodeepseek-v4-flash
- 产品
- —