字节跳动发布 Seedance 1.0

短视频帝国把视频生成写成云与豆包入口

火山引擎 Force 大会发布 Seedance 1.0 与 Pro 等视频生成基座,支持文/图输入与多镜头叙事,并接入豆包等内容入口,把字节的视频生成能力推到公开产品前台。

时间2025 年 6 月 11 日 级别A · 行业级 组织字节跳动 / ByteDance 状态已核验 · 1 个来源
编辑插图:分镜格子连成一条流动的舞蹈光轨
AI Chronicle 原创插图:分镜连成 Seedance(种子之舞)的运动轨迹。 AI Chronicle

Seedance 1.0 技术报告里有一条很像分镜脚本的提示:侦探走进昏暗房间,检查桌上的线索,拿起一件物品,镜头再切到他思考。过去的视频模型往往要求创作者把这些动作拆开生成,再靠剪辑维持人物和空间。2025 年 6 月 11 日,字节在火山引擎 Force 大会发布 Seedance 1.0 时,把“原生多镜头”放到了模型名字旁边。

官方描述的典型输出是十秒视频,其中包含两到三次镜头转换,可以在远景、中景与特写之间切换。模型统一支持文生视频和图生视频,目标分辨率为 1080p。技术报告还给出一个速度样本:在 NVIDIA L20 上生成五秒 1080p 视频约需 41.4 秒。这个数字说明字节优化了什么,也说明结果并非即时出现;换硬件、提示和服务负载后,等待时间会变化。

多镜头并不只是把三个短片接在一起。前一镜出现的人物、道具和动作意图,要在下一镜继续成立。为此,Seed 团队描述了镜头边界检测、自动视频字幕、统一的多任务训练与面向视频的反馈优化。Artificial Analysis 当时将它排在文生视频和图生视频榜单前列,但排行榜只能反映特定样本和投票环境。创作者最后仍要看自己的提示需要重试多少次。

Seedance 1.0 的另一半发生在模型之外。字节拥有豆包、即梦、剪映生态和火山引擎 API,可以把同一套生成能力分别交给普通用户、专业创作者和企业开发者。对一家短视频平台公司来说,分发不是发布后的附加项,而是基座设计的一部分:生成结果会在哪里被修改、传播,最终又怎样回到内容平台。

这次发布让字节第一次有了一个外界可以稳定点名的视频模型品牌。此前,人们容易把它的视频 AI 理解成剪辑工具里的特效或推荐系统背后的能力;Seedance 把底层模型、技术报告和产品入口连成一条线。十秒片段仍然短,多镜头仍会漂移,但字节从此不再只是视频生成的渠道,也成为必须被直接检验的模型供应者。

One prompt in the Seedance 1.0 technical report reads like a compact storyboard. A detective enters a dim room, examines clues on a table, picks up an item, and the camera changes to capture him thinking. Earlier video workflows often asked a creator to generate those actions separately and use editing to preserve the person and the space. When ByteDance released Seedance 1.0 at the Volcano Engine Force conference on June 11, 2025, it placed “native multi-shot” beside the model’s name.

The representative output described by the company was a ten-second video containing two or three shot transitions, moving among wide, medium, and close views. One model handled both text-to-video and image-to-video work at a target resolution of 1080p. The report also supplied a concrete speed sample: on an NVIDIA L20, a five-second 1080p generation took about 41.4 seconds. That figure showed where ByteDance had invested in acceleration and where the boundary remained. Generation was not instantaneous, and different hardware, prompts, and service load would change the wait.

Multi-shot generation required more than placing three short clips together. A person, object, and intended action established in one view had to remain meaningful in the next. The Seed team described shot-boundary detection, automatic video captioning, unified multitask training, and feedback optimization adapted to video. Artificial Analysis ranked the model highly in text-to-video and image-to-video at the time, but a leaderboard reflected a particular sample and voting environment. A creator’s practical result still depended on how many retries a private prompt consumed.

The other half of Seedance 1.0 existed outside the model. ByteDance could place related generation capabilities in Doubao, Jimeng, its editing ecosystem, and Volcano Engine APIs, reaching consumers, professional creators, and application developers through different doors. For a short-video platform company, distribution was not an accessory added after research. It shaped where an output would be revised, shared, and eventually returned to a content network.

The release gave ByteDance a stable video-model name that outsiders could cite and test. Its video AI had previously been easy to perceive as an effect inside an editing tool or a capability hidden behind recommendation systems. Seedance connected a foundation model, a technical report, and product distribution. Ten seconds remained short and shot-to-shot continuity could still fail, but ByteDance was no longer only a channel for generated video. It had become a model supplier responsible for what appeared in the frame.

展开完整事件档案人物、主题、模型与产品
人物
模型
seedance
产品
doubaovolcengine-ark
来源

原始资料

  1. 01Seedance 1.0 tech report noteByteDance Seed · official

试试搜索