艺术家集体起诉生成式 AI 公司
训练数据的版权问题被推上法庭
一批艺术家对 Stability AI、Midjourney 等公司提起集体诉讼,指控其用受版权保护的图像训练模型。这开启了对 AI 训练数据合法性的全球法律博弈。
2023 年初,一批艺术家对 Stability AI、Midjourney 等公司提起集体诉讼,核心指控只有一句话:你们用我们的作品训练了模型。这些模型能生成特定画家的风格图像,而画家本人从未授权、也未获得分成。这起诉讼的象征意义,远大于它具体的法律诉求——它第一次把「AI 用了谁的数据、该不该付钱」推上了法庭。
这场争议的根源,在生成式 AI 诞生之初就埋下了。Stable Diffusion、Midjourney 这些模型之所以强大,是因为它们吞下了互联网上海量的图像与文本——其中大量是受版权保护的作品。在模型还只是实验室玩具时,这种抓取无人问津;但当生成能力开始冲击艺术家的饭碗,问题就再也无法回避。
诉讼带来的第一个变化,是行业对训练数据的重新审视。AI 公司开始强调数据合规、建立授权机制、参与内容溯源标准;平台开始要求标注 AI 生成内容;创作者则开始为自己的作品进入训练集争取话语权。这些今天被视为「行业标配」的做法,起点都可以追溯到 2023 年的这批诉讼。
当然,版权争议远比一句「剽窃」复杂。模型到底是在「复制」还是在「学习」、风格算不算受保护的对象、合理使用边界在哪里——这些问题的答案至今仍在司法与政策的博弈中演化。诉讼的结果有输有赢,但输赢本身或许不如「问题被公开讨论」更重要。
回看艺术家集体起诉生成式 AI 公司,它像一场迟到的正式见面礼:当 AI 的能力开始触碰人的创作与生计,围绕数据、版权与利益的规则就必须被建立起来。2026 年欧盟 AI 法案对训练数据的透明要求、各大平台的授权协议,都是这场 2023 年初开启的博弈的延续。
In early 2023 a group of artists filed class-action suits against Stability AI, Midjourney, and others, with one core accusation: you trained your models on our works. These models can generate images in a specific painter's style, while the painter never authorized it and never received a share. The symbolic weight of these suits far exceeded their specific legal demands—they pushed "whose data did AI use and should it pay" before the courts for the first time.
The roots of this dispute were buried at the birth of generative AI. Models like Stable Diffusion and Midjourney became powerful because they ingested vast amounts of images and text from the internet—much of it copyrighted works. When models were lab toys, nobody cared about the scraping; but once generation began threatening artists' livelihoods, the question could no longer be avoided.
The first change the suits drove was a re-examination of training data across the industry. AI companies began emphasizing data compliance, building licensing mechanisms, and joining content-provenance standards; platforms started requiring AI-content labels; creators began demanding a voice in whether their works enter training sets. Practices today treated as "industry standard" can all trace their start to the 2023 suits.
Of course, the copyright dispute is far more complex than "plagiarism." Whether a model is "copying" or "learning," whether style counts as protectable subject matter, and where fair use ends—these questions are still evolving through judicial and policy struggles. The suits had winners and losers, but perhaps the outcome matters less than the fact that the question was publicly aired.
Looking back at the artists' mass suit against generative-AI companies, it reads like a belated formal introduction: once AI's capabilities touch human creation and livelihood, rules about data, copyright, and interests must be built. The EU AI Act's transparency requirements for training data in 2026 and platforms' licensing agreements are all continuations of the battle that opened in early 2023.
展开完整事件档案人物、主题、模型与产品
- 人物
- —
- 模型
- —
- 产品
- —