Midjourney 发布 V6

文生图开始认真对待提示里的字

Midjourney V6 在提示遵循与文字渲染上明显进步,使社区默认美学从“氛围滤镜”更接近可执行的设计说明。

时间2023 年 12 月 21 日 级别B · 领域级 组织Midjourney 状态已核验 · 1 个来源
编辑插图:画框中隐约浮现清晰字母轮廓,周围是柔和的绘画笔触
AI Chronicle 原创插图:画框里的字开始可读,对应 V6 对提示与文字渲染的推进。 AI Chronicle

2023 年 12 月 21 日,Midjourney 用户在 Discord 里把模型切到 V6 alpha,随后做起了最朴素的实验:复制一条在 V5 上用过的提示,再生成四张。社区不需要等一份完整评测报告;两组图片并排摆出来,衣服颜色、人物位置、镜头关系有没有被听见,一眼就能开始争论。这场集体盲测没有裁判,却比任何榜单都更快地传开了结论。

最受关注的变化之一,是画面里的短文本。过去让模型做一张带标题的海报,字母通常只是像文字的纹理,设计师只能留白后再去排版软件补字。V6 仍会拼错、漏字,也经不起长句,但被引号标出的短词更常呈现出可读轮廓。这个进步没有让平面设计自动完成,却让包装草图、招牌和概念海报多了一种以前很难测试的可能。对电商和广告团队,这意味着提案阶段可以少一轮“先找字体”的等待。

更广的变化发生在提示本身。V5 时代形成了大量咒语式写法:堆镜头名、风格词、材质词,希望模型从词云中抓住气氛。V6 更愿意处理句子里的关系与细节,用户也被迫重新学习——少一点无意义的修饰,多说清楚谁在什么位置、穿什么、画面上需要出现什么。提示词开始接近一份粗糙的创意简报,而不是一串召唤咒语。

这种能力通过一个很特殊的产品形态扩散。Midjourney 没有开放权重,V6 初期也不是独立设计软件里的正式按钮,而是订阅用户在 Discord 中输入参数选择的 alpha 模型。模型、教程和审美反馈挤在同一个社区里:有人发失败图,有人马上改写提示,新的用法在几个小时内被复制。产品的不便与它的传播速度同时存在——没有应用商店,却有最密集的现场教学。

V6 的意义不在于从此每个字都能正确生成,而在于文生图的评价标准发生了变化。图片“看起来很好”已经不够,模型还要对要求负责。废图没有消失,设计师也没有退出流程;只是筛图时多了一个更严格的问题:这张图除了漂亮,是否真的完成了我写下的那件事?

On December 21, 2023, Midjourney users switched their Discord jobs to the V6 alpha and began the simplest possible experiment: copy a prompt previously used with V5 and generate another grid. The community did not need to wait for a formal evaluation suite. Put the images beside each other and arguments could begin immediately. Did the clothing color survive? Were two people placed where the sentence requested? Did the camera relationship still make sense?

Short text inside the image became one of the most visible tests. Earlier versions could produce a beautiful poster whose title was only a texture shaped like language. Designers learned to leave space and add real type later in another application. V6 still misspelled words, dropped letters, and struggled with longer phrases, but short quoted text more often appeared with readable structure. It did not automate graphic design. It made signs, packaging studies, and concept posters possible to explore in a way that had previously failed at the first word.

A broader change occurred in the prompt itself. V5 had produced an incantation culture: users stacked camera names, style labels, materials, and quality terms in the hope that the model would recover the intended atmosphere. V6 responded more often to relationships and concrete details inside a sentence. Users had to relearn their craft—remove decorative prompt debris and state who was where, what they wore, and which element needed to appear in the frame. A prompt began to resemble a rough creative brief.

The capability spread through an unusual product form. Midjourney did not release model weights, and V6 initially was not a polished button inside a standalone design suite. It was an alpha selected through parameters by subscribers in Discord. Model access, tutorials, and aesthetic feedback occupied the same community. Someone posted a failure, another person adjusted the wording, and a new convention could be copied within hours. The friction of the interface and the speed of collective learning were two sides of the same distribution system.

V6 mattered not because every generated word became correct, but because the standard for consumer text-to-image systems changed. “It looks good” was no longer sufficient; the image increasingly had to answer to the request. Rejects did not disappear, and designers did not leave the loop. Selection simply gained a harder question: beyond being attractive, did this image actually do the thing the brief asked for?

展开完整事件档案人物、主题、模型与产品
人物
模型
midjourney-v6
产品
来源

原始资料

  1. 01MidjourneyMidjourney · official

试试搜索