Meta 发布 LLaMA
高性能权重离开 API,开放模型社区迅速成形
Meta 发布 7B、13B、33B 和 65B 参数的 LLaMA 论文,并向获批研究者提供非商业许可权重;其中 65B 模型用 1.4 万亿 token 训练。权重随后外泄,社区迅速开发量化、微调和本地推理工具。
2023 年 2 月 24 日,Meta 发表 LLaMA 论文,并公布 7B、13B、33B、65B 四档规模;65B 用约 1.4 万亿 token 训练。权重以非商业研究许可向获批研究者提供——不是 OSI 意义上的开源软件,也不能直接写进大多数产品合同。申请、审批与研究用途,是官方分发的门槛。
论文强调:用更多 token 训练相对较小的模型,可在多项基准上以较少参数达到竞争表现。对仍被封闭 API 挡住的研究者,这句话的吸引力很具体:检查权重、改训练、在自己的硬件边界内跑模型,忽然显得可能。小模型加长数据的路线,也给算力不那么充裕的实验室留下缝隙。
随后权重外泄。外泄不是官方发布策略,却迅速改变了事实地图:社区连夜做量化、微调与本地推理工具,LLaMA 派生模型与教程在几周内铺开。能力讨论与合规讨论被迫并行——能跑不等于被允许商用。有人把外泄当成“开放的胜利”,许可文本仍写着非商业;两种说法同时成立,也同时不完整。
LLaMA 留下的分界比任何单次分数更清晰:开放权重、开放代码、开源许可是三件不同的事。第一代把前两件中的“权重可获得”推到台前,第三件要等到 Llama 2 才以另一套社区许可部分回应。外泄加速了工具链,也加速了行业学习如何把法律条款写进技术故事。后来所有“开源大模型”争论,几乎都要回到这一天画下的三条线。
外泄之后的几个星期,Hugging Face 与各类脚本把量化、GGUF、LoRA 与本地 UI 推到普通人电脑上。研究许可挡不住传播;传播也改不了许可原文。Meta 后来用 Llama 2 的社区许可回应商业需求,但第一代留下的教训已经写进行业词典:谈“开源模型”时,先问开放的是权重、代码,还是许可证。三条线画不清,争论只会变成口号对轰。
非商业许可下的研究分发,本意是控制用途;外泄把控制权部分转移给社区事实。此后任何“研究专用权重”策略,都要同时设计防泄漏与泄漏后的叙事。LLaMA 不是第一个外泄案例,却是大模型时代最有示范效应的一次。
On 24 February 2023, Meta published the LLaMA paper and four sizes—7B, 13B, 33B, 65B—with the 65B model trained on about 1.4 trillion tokens. Weights went to approved researchers under a noncommercial research license: not open-source software in the OSI sense, and not something most product contracts could absorb directly. Application, approval, and research use were the official distribution gates.
The paper stressed training relatively smaller models on more tokens to reach competitive benchmarks with fewer parameters. For researchers still blocked by closed APIs, the appeal was concrete: inspect weights, change training, run inside one’s own hardware boundary—suddenly plausible. Smaller models with longer data also left a gap for labs with less compute.
Then the weights leaked. The leak was not official release strategy, yet it quickly rewrote the map of facts: the community built quantization, fine-tuning, and local inference tools overnight; LLaMA derivatives and tutorials spread in weeks. Capability talk and compliance talk were forced to run in parallel—being able to run is not being allowed to ship. Some called the leak a victory for openness; the license still said noncommercial. Both statements can be true and both incomplete.
LLaMA left a clearer boundary than any single score: open weights, open code, and open-source licenses are three different things. The first generation pushed “weights obtainable” of the first two to the front; the third waited for Llama 2’s different community license to answer in part. The leak sped the toolchain and sped the industry’s education in writing legal terms into technical stories. Later debates over “open-source large models” almost always return to the three lines drawn this day.
In the weeks after the leak, Hugging Face and assorted scripts pushed quantization, GGUF, LoRA, and local UIs onto ordinary machines. A research license could not stop spread; spread could not rewrite the license text. Meta later answered commercial needs with Llama 2’s community license, but the first generation’s lesson was already in the industry lexicon: when speaking of “open-source models,” ask first whether weights, code, or the license is open. Until those three lines are clear, debate collapses into slogan combat.
Research distribution under a noncommercial license meant to control use; the leak partly moved control to community facts. Any later “research-only weights” strategy must design both against leaks and for post-leak narrative. LLaMA was not the first leak, but it was among the most exemplary in the large-model era.
展开完整事件档案人物、主题、模型与产品
- 人物
- —
- 模型
- llama
- 产品
- —