DeepSeek-R1 发布
开放推理模型引发全球重新定价
DeepSeek 发布开放推理模型 DeepSeek-R1,以强化学习驱动的推理能力和开放权重迅速引发全球关注。
DeepSeek-R1 技术报告里最耐读的,往往不是最高那一格分数。R1-Zero 与 R1 两个名字之间,留着一段不太体面的距离。成功叙事若删掉这段距离,就只剩广告。
2025 年 1 月,DeepSeek 发布开放权重推理模型 R1 与 R1-Zero,并附六个蒸馏小模型。R1-Zero 被直接放进大规模强化学习,前面没有常规监督微调示范“好答案该怎么写”。奖励来自可核验结果的任务;模型在尝试中长出更长推理、自检与改策略。报告把这些视为纯 RL 能激发推理的证据,也立刻承认:它会无休止重复,文字难读,还会语言混杂。会做题与会把思路交给人读,不是同一件事。
R1 没有把 Zero 的野生状态奉为理想。开发者加入冷启动数据,再安排多阶段强化学习与监督微调。它仍继承 DeepSeek-V3-Base 的 MoE:总参数 6710 亿,每 token 激活 370 亿。官方表中 R1 在 AIME 2024 pass@1 为 79.8%,MATH-500 为 97.3%,LiveCodeBench 为 65.9%;SimpleQA 等项仍明显低于 o1。数字来自发布方设定;采样温度与 pass@1 估计方式应一并留下。
另一条线是蒸馏:用 R1 推理样本微调 Qwen 与 Llama 系列,开放 1.5B 到 70B 六档稠密模型。仓库还保留不英雄的使用建议:温度别太低,避免无尽重复;有些题需显式要求以思考标记开始。榜单上的强,到运行仍靠细小条件维持。可验证奖励主要来自数学与代码,解释了为何 STEM 表刺眼而开放事实未必同样耀眼。
“突然会推理”的神话容不下这份报告的粗糙。路线更曲折:先在可验证奖惩中长出行为,再承认难读,再整理成人能用的东西,最后压进更小模型。幕与幕之间的缺陷不被删去——在分数会被迅速追平的领域,这比任何单次第一更耐读。
蒸馏六模型让“推理能力”不再只属于 671B 级检查点。小模型能否在可接受成本下复现部分行为,成为教育、竞赛与本地部署真正关心的问题。报告把缺陷留在正文里,等于邀请读者用实验态度而非神话态度接近 R1。分数会旧,这种写法更耐放。
In the DeepSeek-R1 technical report, the most durable pages are often not the top score cell. Between the names R1-Zero and R1 sits an unflattering distance. Delete that distance from the success story and only advertising remains.
In January 2025 DeepSeek released open-weight reasoning models R1 and R1-Zero, plus six distilled smaller models. R1-Zero went straight into large-scale reinforcement learning without the usual supervised fine-tuning that demonstrates how a “good answer” should read. Rewards came from checkable tasks; through trial the model grew longer reasoning, self-checks, and strategy switches. The report treats those behaviors as evidence that pure RL can elicit reasoning—and immediately admits endless repetition, hard-to-read prose, and mixed languages. Solving problems and handing a readable chain of thought to a person are not the same skill.
R1 does not canonize Zero’s wild state. Developers add cold-start data, then multi-stage RL and supervised fine-tuning. It still inherits DeepSeek-V3-Base’s MoE: 671 billion total parameters, 37 billion activated per token. On official tables R1 reaches 79.8% AIME 2024 pass@1, 97.3% MATH-500, 65.9% LiveCodeBench; on SimpleQA and similar it remains clearly below o1. Numbers are the publisher’s settings; temperature and pass@1 estimation should travel with the scores.
Another line is distillation: R1 reasoning samples fine-tune Qwen and Llama series into six dense models from 1.5B to 70B. The repository keeps unheroic usage notes—do not set temperature too low; some problems need an explicit request to begin with thinking markers. Strength on a leaderboard still depends on small runtime conditions. Verifiable rewards mainly from math and code help explain why STEM tables glare while open-domain fact QA may not.
Myths of “suddenly reasoning” cannot hold this report’s roughness. The route is more crooked: grow behavior under verifiable reward, admit unreadability, shape something people can use, then press capability into smaller models. Defects between acts are not deleted—in a field where scores are soon matched, that is more durable than any single first place.
Six distilled models mean “reasoning ability” no longer belongs only to a 671B-class checkpoint. Whether small models can reproduce part of the behavior at acceptable cost becomes what education, contests, and local deploy really care about. Leaving defects in the main text invites readers to approach R1 experimentally rather than mythically. Scores age; that writing style keeps.
展开完整事件档案人物、主题、模型与产品
- 人物
- —
- 模型
- deepseek-r1
- 产品
- —