DALL·E 2 让文生图进入实用阶段

自然语言提示变成高质量图像的转折点

OpenAI 发布 DALL·E 2,用扩散模型把文本描述转成高分辨率图像,支持编辑与变体生成。它让「文生图」从研究演示变成大众可体验的创作工具。

时间2022 年 4 月 6 日 级别A · 行业级 组织OpenAI 状态已核验 · 1 个来源
发光画笔绘制超现实景观的插画
DALL·E 2 把文本生成图像带进主流视野,让创意工作第一次大规模用上生成式 AI。 AI Chronicle

2022 年 4 月,OpenAI 发布了 DALL·E 2。它的前一代 DALL·E 在 2021 年已经证明「文本生成图像」可行,但画质有限,更像研究演示。DALL·E 2 把这件事往前推了一大步:输入一句「一只宇航员骑马的油画」,几秒钟后得到一张高分辨率、语义精准、风格可控的图像。文生图从此进入可用的阶段。

技术上,DALL·E 2 是组合了 CLIP 与扩散模型的系统。CLIP 负责把文本和图像映射到同一个语义空间,让模型理解「一句话对应什么样的画面」;扩散模型负责从噪声中一步步还原出符合描述的图像。它还支持对已有图像做局部编辑、生成不同变体。这套「语义理解+像素生成」的分工,成为后来整个文生图领域的技术底色。

发布的方式也很有 2022 年特色:候补名单。用户提交申请,等待审核通过才能使用。这种「先到先得」的稀缺感反而放大了热度——社交媒体上,第一批用户生成的图像被疯狂转发,「提示词工程」这个概念开始进入大众词汇。人们第一次意识到:跟 AI 描述得越清楚,得到的图就越好。

DALL·E 2 引爆的不只是热度,还有争议。艺术家担忧自己的风格被模仿,平台担忧版权归属,还有人担忧深度伪造与误导性图像。OpenAI 当时回应说,模型拒绝生成名人面孔和暴力内容,并给生成图像打上水印。这些讨论后来演变成全球范围内关于 AI 生成内容监管的持久议题——而起点,就在 2022 年的这次发布。

DALL·E 2 的产业影响迅速扩散。同年 7 月 Midjourney 在 Discord 走红,8 月 Stable Diffusion 开源引爆全民创作。短短几个月内,文生图从少数公司的实验变成了一场大众运动。OpenAI 自己则在 2023 年推出 DALL·E 3,把生成能力直接集成进 ChatGPT——提示词创作彻底成为产品的一部分。

回看 2022 年 4 月,DALL·E 2 最持久的意义,是把「用语言创作图像」从可能性变成了现实,并且让所有人都能触摸到。它定义了一种新的创作媒介——提示词——也提前暴露了这种媒介带来的所有问题:版权、真实、伦理。技术后来被无数产品超越,但那个「一句话生成一张图」的时刻,是生成式 AI 大众化叙事里绕不开的原点。

In April 2022 OpenAI released DALL·E 2. Its predecessor DALL·E had proven in 2021 that text-to-image was possible, but image quality was limited and it read like a research demo. DALL·E 2 pushed things far forward: type "an astronaut riding a horse in oil painting style," and within seconds you get a high-resolution, semantically precise, stylistically controllable image. Text-to-image had entered a usable stage.

Technically, DALL·E 2 combined CLIP with a diffusion model. CLIP mapped text and images into the same semantic space, letting the model understand "what picture does this sentence correspond to"; the diffusion model progressively restored an image from noise to match the description. It also supported local editing of existing images and generating variations. This division of "semantic understanding plus pixel generation" became the technical foundation of the whole field.

The launch style was also very 2022: a waitlist. Users applied and waited for approval. The scarcity amplified the buzz—the first users' images were retweeted wildly, and "prompt engineering" entered the popular vocabulary. People realized for the first time: the clearer you describe to the AI, the better the image.

DALL·E 2 ignited not just hype but controversy. Artists worried their styles would be imitated, platforms worried about copyright, and others worried about deepfakes and misleading images. OpenAI responded that the model refused celebrity faces and violent content and watermarked outputs. These debates grew into a lasting global issue about AI-generated content regulation—and their starting point was this release.

The industry impact spread quickly. In July Midjourney went viral on Discord; in August Stable Diffusion's open release ignited mass creation. Within months text-to-image went from a few companies' experiments to a mass movement. OpenAI itself shipped DALL·E 3 in 2023, integrating generation directly into ChatGPT—prompt-based creation had fully become part of the product.

Looking back at April 2022, DALL·E 2's most durable significance was turning "creating images with language" from possibility into reality, touchable by everyone. It defined a new creative medium—the prompt—and exposed early all the problems that medium brings: copyright, truth, ethics. Later technology overtook it, but that moment of "one sentence, one image" remains the point of origin in the narrative of generative AI's democratization.

展开完整事件档案人物、主题、模型与产品
人物
模型
产品
来源

原始资料

  1. 01DALL·E 2OpenAI · official

试试搜索