Google 发布 Veo 3 原生音频视频

声画同步成为视频模型的标配野心

Google 在 I/O 2025 周期发布 Veo 3,强调原生同步音频(对白、音效、环境声)与更强可控性,并逐步接入 Gemini 应用与 API。

时间2025 年 5 月 20 日 级别A · 行业级 组织GoogleGoogle DeepMind 状态已核验 · 1 个来源
编辑插图:波形嵌入电影画幅
AI Chronicle 原创插图:画幅里的声波,对应 Veo 3 原生音频。 AI Chronicle

2025 年 5 月 20 日,Google 在 I/O 上发布 Veo 3,并把它放进新的 AI 电影制作工具 Flow。发布示例的提示不只描述画面,还写进“树叶摩擦”“发动机轰鸣”“角色对话”等声音要求。模型尝试在生成镜头的同时生成对白、音效与环境声,视频提示因此从一段视觉描述变成一页带声场的微型剧本。

在此前的大多数公开文生视频工作流中,声音属于下一步。创作者先挑选可用画面,再去配乐、拟音、录对白并调整节奏。Veo 3 将这些轨道提前合并:角色说话时口型与声音需要相互配合,物体撞击要在正确时间发声,镜头从室内移到室外时环境声也要变化。生成的对象不再只是连续图像,而是一段需要同时维持视觉与听觉因果的时间结构。

这项能力通过具体的产品门槛进入现实。Flow 首发面向美国的 Google AI Pro 与 Ultra 订阅用户;Pro 每月提供 100 次生成,Ultra 给予更高限额,并提供 Veo 3 原生音频的早期访问。Gemini 与开发者入口随后逐步扩展。发布会上的连续演示没有排队、配额和地区问题,普通用户的体验却由这些条件共同决定。

原生音频也不会自动省掉后期。生成对白可能不够清楚,口型、语气与角色身份可能漂移,环境声可能在镜头切换时断裂。对创作者而言,一次性得到声画草稿可以更快判断场景是否成立,却也把错误绑得更紧:过去可以单独替换一条音轨,现在可能需要重新生成整段,或把模型输出拆回传统时间线修理。

风险也随声音一起增加。无声的虚假画面已经能误导,有同步对白与环境声的片段更容易被当作现场记录。Google 为生成内容使用 SynthID 等来源标识与安全措施,但水印的可检测性、平台是否保留标记、恶意剪辑能否绕过识别,仍不是模型本身可以解决的问题。声音提高了表现力,也提高了伪造的说服力。

Veo 3 首发真正改变的,不是给视频模型多加一个功能复选框,而是改变了“生成一个镜头”的最小单位。画面、台词、拟音和氛围第一次被要求在同一次推理中对齐。对于电影制作,它提供了一种更完整也更难修改的草稿;对于评测,漂亮画面已经不够,声音是否在正确时刻属于正确对象,也成为模型必须通过的考试。

Google announced Veo 3 at I/O on 20 May 2025 and placed it inside a new AI filmmaking tool called Flow. The launch prompts described more than images. They asked for rustling leaves, roaring engines, character dialogue, and environmental sound. The model attempted to create speech, effects, and ambience with the footage itself. A video prompt was no longer only a visual paragraph; it had begun to resemble a miniature screenplay with a sound field.

In most earlier public text-to-video workflows, audio belonged to the next stage. A creator selected usable footage, then added music, foley, dialogue, and timing in an editor. Veo 3 moved those tracks forward. A speaking character required voice and mouth movement to correspond. An impact needed to sound at the correct moment. When a camera moved from indoors to outdoors, the ambience had to change with it. The generated object was not merely a sequence of images but a temporal structure expected to preserve visual and auditory cause together.

The capability reached users through specific product gates. Flow launched in the United States for Google AI Pro and Ultra subscribers. Pro offered 100 generations per month; Ultra provided higher limits and early access to Veo 3 with native audio. Gemini and developer channels expanded later. A continuous keynote reel contained no queue, region, or quota. Ordinary use was shaped by all three.

Native sound did not automatically eliminate post-production. Dialogue could be unclear, lip movement and speaker identity could drift, and ambience could break across a cut. Receiving an audio-visual draft in one pass let creators judge a scene more quickly, but also bound its errors together. Previously, an editor could replace one audio track. Now a failure might require regenerating the entire clip or breaking the model’s output back into a conventional timeline for repair.

Risk increased with sound as well. Silent synthetic footage could already mislead; synchronized dialogue and ambience made a clip easier to mistake for a record of an actual event. Google attached provenance and safety measures including SynthID to generated media, but watermark detection, platform preservation, and malicious re-editing were not problems the model alone could solve. Audio increased expressive power and the persuasive force of a forgery.

The consequential change in Veo 3 was not an additional checkbox on a video model. It altered the minimum unit meant by “generate a shot.” Image, line reading, foley, and atmosphere were asked to align in the same inference. For filmmaking, this produced a more complete draft that could also be harder to revise. For evaluation, attractive frames were no longer sufficient: the model also had to make the right sound happen at the right time to the right object.

展开完整事件档案人物、主题、模型与产品
人物
模型
veo-3
产品
gemini
来源

原始资料

  1. 01VeoGoogle DeepMind · official

试试搜索