DeepSeek-V2 发布

MLA 架构与超低价 API 引爆中国大模型价格战

DeepSeek 发布 V2,采用 MLA(多头潜在注意力)+ DeepSeekMoE 稀疏架构,以远低于同行的价格开放 API。推理成本约为 Llama 3 的 1/100,直接引爆了中国大模型的价格战。

时间2024 年 5 月 6 日 级别A · 行业级 组织DeepSeek(深度求索) 状态已核验 · 1 个来源
蓝色鲸鱼游弋在数据电路海洋中的插画
DeepSeek-V2 用 MLA 与 MoE 架构引爆了 2024 年的模型价格战。 AI Chronicle

2024 年 5 月的中国大模型市场,正处于最热闹也最拥挤的时刻:几乎每周都有新模型发布,各家都在标榜自己更强。在这片喧嚣中,DeepSeek 扔出了一张出人意料的牌——不是单纯拼能力,而是把价格打到了地板以下。

5 月 6 日,DeepSeek-V2 发布。它的技术底子相当硬:MLA(多头潜在注意力)架构大幅压缩了推理时的内存开销,配合稀疏 MoE,让效率远超同级模型。在多项中文基准上,V2 追平甚至超过了 Llama 3 70B。但真正引爆行业的,是它的定价——每百万 token 只要 1 元,大约是主流闭源模型的百分之一。

价格战立刻开打。V2 发布后几天内,字节、阿里、百度、腾讯、科大讯飞等纷纷跟进降价,有厂商甚至宣布部分模型免费。行业从「拼参数」转向「拼成本」,推理效率第一次成为公开叫板的卖点。对开发者来说,这意味着调用大模型的门槛从「贵」变成了「随便用」。

低价背后不是慈善,是工程能力的变现。MLA 和稀疏架构让同样一次推理的算力成本大幅下降,DeepSeek 才有底气把价格砍到别人不敢跟的地步。技术上的领先,最终以最刺眼的方式——价格——呈现在市场上。

回看 V2,它是「技术突破引发商业地震」的典型案例。它没有像后来 R1 那样制造全球性轰动,但它做了一件更基础的事:让所有人重新评估「用大模型到底该花多少钱」。当整个行业的价格体系被重新锚定,当推理效率成为竞争的核心,DeepSeek 已经为它后来的传奇故事,打下了第一块最坚实的基石。

In May 2024, China's large-model market was at its liveliest and most crowded moment: new models launched almost weekly, each claiming to be stronger. In that noise, DeepSeek played an unexpected card—not pure capability, but prices driven to the floor.

On May 6, DeepSeek-V2 was released. Its technical foundation was solid: the MLA (Multi-head Latent Attention) architecture drastically cut memory overhead at inference, and with sparse MoE, efficiency far exceeded comparable models. On many Chinese benchmarks, V2 matched or surpassed Llama 3 70B. But what truly detonated the industry was pricing—1 yuan per million tokens, roughly one-hundredth of mainstream closed models.

The price war broke out immediately. Within days of V2, ByteDance, Alibaba, Baidu, Tencent, iFlytek and others followed with cuts; some vendors even declared parts of their offerings free. The industry shifted from "compare parameters" to "compare costs," and inference efficiency became a publicly contested selling point for the first time. For developers, the threshold for calling a large model went from "expensive" to "use it freely."

Behind the low price was not charity but the monetization of engineering. MLA and sparse architecture cut the compute cost of a single inference dramatically, giving DeepSeek the confidence to slash prices where others dared not follow. Technical leadership was finally presented to the market in the most eye-catching form: price.

Looking back, V2 is a classic case of "technical breakthrough triggers commercial earthquake." It did not create the global sensation R1 later did, but it did something more foundational: it made everyone reassess how much using a large model should cost. As the industry's price system was re-anchored and inference efficiency became the core of competition, DeepSeek had laid the first and sturdiest cornerstone of its later legend.

展开完整事件档案人物、主题、模型与产品
人物
模型
产品
来源

原始资料

  1. 01DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language ModelDeepSeek / arXiv · paper

试试搜索