Claude 3.5 Sonnet 发布
中档模型开始承担旗舰工作
Anthropic 发布 Claude 3.5 Sonnet,以更低的价格和更快的速度达到接近或超过上一代旗舰的代码、视觉和推理表现。
2024 年夏天,一个接入 Claude API 的团队若要选默认模型,通常会先做一道取舍题:最强的那档太贵、太慢,便宜的那档又未必扛得住复杂代码和长文档。6 月 20 日,Anthropic 把 Claude 3.5 Sonnet 放进了这道题的中间位置,却给了它接近旗舰的任务清单。
数字让这次换位变得具体。Claude 3.5 Sonnet 的 API 价格是每百万输入 token 3 美元、输出 15 美元,与 Claude 3 Sonnet 同档;上下文窗口为 20 万 token。Anthropic 还称它的运行速度达到 Claude 3 Opus 的两倍。代码、视觉推理和复杂指令上的成绩来自厂商评测,不能直接外推到所有工作负载,但对每天要计算延迟和账单的开发者来说,价格与速度不是榜单上的抽象分数。
同一天出现的 Artifacts,把这种能力变化从 API 表格带到了屏幕上。用户让 Claude 写网页、代码片段或文档时,结果不再只顺着聊天记录往下堆,而是可以出现在对话旁边的独立窗口里,继续查看、修改和迭代。它当时仍是预览功能,离稳定的协作环境还有距离;可界面已经给出一个新隐喻:模型产出的不只是回复,也可以是正在施工的东西。
这两条线很快汇合。更快、价格可控的 Sonnet 适合被开发工具反复调用;Artifacts 又让普通用户看到“生成—修改—再生成”的工作循环。旗舰模型过去负责展示天花板,3.5 Sonnet 则争夺每天被点击最多的那个位置。对生产系统而言,默认值往往比冠军更有力量:教程会围绕它写,脚手架会预填它,团队会以它的成本建立预算。
Claude 3.5 Sonnet 并没有证明中档模型永远胜过旗舰。它完成的是一次更实际的换位:模型家族不再只按能力从小到大排队,而开始按工作负载分工。此后,“该用哪一个模型”越来越像路由问题——只有少数任务需要冲向最贵的一档,大量日常工作留在那个足够聪明、足够快,也足够便宜的中间层。
In the summer of 2024, a team choosing a default Claude model faced a familiar compromise. The most capable tier could be too expensive or slow to call on every request; the cheaper tier might not be dependable enough for a large code change or a long document. On June 20, Anthropic placed Claude 3.5 Sonnet in the middle of that price ladder while giving it the job description of a flagship.
The numbers made the repositioning tangible. Claude 3.5 Sonnet cost $3 per million input tokens and $15 per million output tokens, unchanged from Claude 3 Sonnet, and offered a 200,000-token context window. Anthropic also said it ran at twice the speed of Claude 3 Opus. Its coding, visual-reasoning, and instruction-following results came from vendor-run evaluations and could not guarantee the same ordering on every private workload. But price and latency were not decorative benchmark columns for developers operating a real service; they determined which model could remain switched on all day.
Artifacts translated the same shift into a product interface. When a Claude.ai user requested a website, a code snippet, or a document, the output could open in a dedicated pane beside the conversation instead of disappearing into the vertical transcript. The feature was still a preview, with the instability and incomplete edges that label implied. Even so, it proposed a different relationship with model output: the generated object was something to inspect and revise, not merely an answer to copy out of a bubble.
The two parts of the launch reinforced each other. A faster, predictably priced Sonnet could be called repeatedly by developer tools; Artifacts made the generate–inspect–revise loop visible to non-API users. Flagship models had traditionally demonstrated the ceiling. Claude 3.5 Sonnet competed for the place clicked most often. In production, that default position can matter more than a benchmark crown: tutorials are written around it, starter code preselects it, and budgets are built from its token price.
Claude 3.5 Sonnet did not establish that a mid-tier model would always outperform a flagship. It established a more durable product idea: model families could be divided by workload rather than arranged as a simple intelligence ladder. After this release, “which model should we use?” increasingly became a routing question. A narrow set of requests could justify the most expensive tier; a much larger stream of everyday work could stay with the model that was capable enough, fast enough, and affordable enough to become routine.
展开完整事件档案人物、主题、模型与产品
- 人物
- —
- 模型
- claude-3.5-sonnet
- 产品
- claude