阿里发布并开源 Qwen3.8-Flash
125B 总参、单 token 激活 6B 的「Next」架构,训练成本较前代骤降约 90%
阿里发布并开源 Qwen3.8-Flash:总参 125B、单 token 激活 6B 的多模态 MoE,采用全新「Next」架构(QSA 稀疏注意力 + Gated DeltaNet 混合注意力),训练成本较 Qwen3.7-Plus 下降约 90%,推理定价每百万输入 1 元、输出 3 元;被视为 Qwen4 的雏形,首发上线「千问办公」。
2026 年 8 月 26 日晚,阿里发布并开源 Qwen3.8-Flash。距离 Qwen3.8-Max 发布只有 23 天。那场发布的主角是 2.4T 总参的旗舰与「下周开源」的承诺;这一场的主角换成了 125B 总参、单 token 激活 6B 的「小」模型——但发布材料里最刺眼的一句话是:训练成本较 Qwen3.7-Plus 下降约 90%。
先看架构。Qwen3.8-Flash 采用全新「Next」架构:QSA(Qwen Sparse Attention)与 GDN(Gated DeltaNet)融合的混合注意力,高缓存命中下 1M 长上下文可提速 8 倍以上;原生支持 262,144 token 上下文,YaRN 可扩展至 1M。官方称 Qwen3.8-Flash-Next 是「Qwen4 的雏形」——这句话比参数表更值得读:阿里把下一代架构先放在一个 6B 激活的模型上验证,而不是等旗舰。
再看成本。训练成本降 90% 需要带着厂商自述的边界读,但推理定价是公开的:每百万输入 1 元、输出 3 元,最低可至 DeepSeek-V4-Flash 的三分之一。这个价格把「便宜到可以随便用」从口号变成现实。当 DeepSeek-V4-Flash 以低价占据开发者心智时,阿里用更低的价格回应——成本曲线成为开源竞争的新标尺。
6B 激活参数同样值得注意。过去「小模型不够强」是默认假设,Qwen3.8-Flash 用 6B 激活在 14 个 benchmark 的 8 个中取得最优结果(官方口径,含 MMLU-Pro、SuperGPQA、SWEBench-Pretrain 等)。这不是说 6B 打败了 2.4T,而是说:对大多数生产任务,6B 激活可能已经够用——「够用」的定义被改写了。
发布节奏也延续了阿里的新打法:模型与产品同场。Qwen3.8-Flash 首发上线「千问办公」,开发者与企业可通过千问 AI 平台获取 API。至此 Qwen3.8 系列已开源 2.4T 的 Max、27B 与 Flash 三大尺寸,全球下载量突破 30 亿次——规模、效率、入口三线并进。
Qwen3.8-Flash 的意义不在单个数字。它把「训练成本降 90%」从口号变成可核验的架构选择,用 6B 激活挑战「小模型不够强」的假设;同时以 Qwen4 雏形的身份预告下一代架构方向。对开发者,这是一个训练与推理成本都极低的 6B 激活多模态模型,可自托管、可微调、可商用;对行业,开源竞争从参数规模转向成本曲线,「训练成本降 90%」成为新的发布叙事。下一场发布的主角是谁,已经不重要——重要的是成本曲线还在往下走。
On the evening of August 26, 2026, Alibaba released and open-sourced Qwen3.8-Flash—23 days after Qwen3.8-Max. That launch starred a 2.4T flagship and a "next week open" promise; this one stars a "small" model with 125B total and 6B active parameters per token. But the sharpest line in the release materials is: training cost about 90% below Qwen3.7-Plus.
On architecture: Qwen3.8-Flash uses the new "Next" architecture—QSA (Qwen Sparse Attention) fused with GDN (Gated DeltaNet) hybrid attention, with 1M long-context speedups above 8x under high cache hits; natively 262,144-token context, extendable to 1M via YaRN. Alibaba calls Qwen3.8-Flash-Next "the Qwen4 prototype"—a line worth more than the parameter table: the next-generation architecture is being validated on a 6B-active model first, not held for a flagship.
On cost: the 90% training-cost claim needs the vendor-self-reported caveat, but inference pricing is public: ¥1 per million input and ¥3 per million output tokens, as low as one-third of DeepSeek-V4-Flash. That price turns "cheap enough to use freely" from slogan into reality. When DeepSeek-V4-Flash captured developers with low prices, Alibaba answered with lower ones—the cost curve became the new yardstick in open competition.
The 6B active parameter count deserves attention too. "Small models are not strong enough" used to be the default assumption; Qwen3.8-Flash reports best results on 8 of 14 benchmarks at 6B active (vendor-reported, including MMLU-Pro, SuperGPQA, and SWEBench-Pretrain). This is not 6B beating 2.4T; it is: for most production tasks, 6B active may already be enough—the definition of "good enough" was rewritten.
The launch rhythm also continues Alibaba's new playbook: model and product on the same stage. Qwen3.8-Flash debuted on Qianwen Office, with API access via the Qwen AI platform. The Qwen3.8 line now spans three sizes—2.4T Max, 27B, and Flash—with global downloads past 3 billion: scale, efficiency, and entry points advancing together.
Qwen3.8-Flash's meaning is not in any single number. It turned "90% lower training cost" from slogan into a verifiable architectural choice, challenging the assumption that 6B active cannot be strong; as the Qwen4 prototype it also previewed the next architecture direction. For developers, this is a multimodal model with 6B active at rock-bottom training and inference cost, self-hostable, fine-tunable, and commercial; for the industry, open competition shifted from parameter scale to cost curves, with "90% cheaper training" becoming the new launch narrative. Who stars in the next launch no longer matters—what matters is that the cost curve is still heading down.
展开完整事件档案人物、主题、模型与产品
- 人物
- —
- 模型
- qwen-3-8-flash
- 产品
- —