GPT-4 发布
图像输入、更稳的复杂任务表现与一次刻意收缩的技术披露
OpenAI 发布 GPT-4,并报告其在模拟律师资格考试中达到约前 10%,而 GPT-3.5 约在后 10%。图像输入最初仅限合作方测试;技术报告未披露参数量、训练数据构成、架构或训练算力。
2023 年 3 月 14 日,OpenAI 发布 GPT-4。技术报告很长。读者可以在里面找到模拟律师资格考试、奥林匹克题目、专业知识测试,也可以看到模型在安全训练前后如何变化。可是最容易记住的,或许是那些找不到的东西:参数规模、训练计算量、数据集构成、具体架构。过去,一篇重要的机器学习论文通常会把机器拆开给同行看;这一次,机器站在评测表后面,只留下运行时投在纸上的影子。
这种阅读经验恰好说明了 GPT-4 所处的位置。它当然还是一个研究对象,却已经不只属于论文共同体。ChatGPT 在几个月里把语言模型推到普通人的桌面,模型能力从实验室指标变成了产品竞争力,也变成了需要严密管理的商业秘密。OpenAI 仍然用技术报告的形式谈论它,但报告的任务已经改变:它不再完整回答“我们造了什么”,而是在证明“它能做到什么”、说明“我们怎样测试风险”,同时划出不会继续公开的边界。
技术报告与产品公告一起构成 3 月 14 日的公共文本。报告里的考试与基准是 OpenAI 的评测叙述;公告里的 Plus 与 API 候补是访问叙述;系统卡里的风险与缓解是安全叙述。三套叙述可以互相引用,却不能互相替代。外部研究者能复现的是 API 行为切片,不能复现的是训练数据配方。这种不对称从此成为前沿模型发布的常态,GPT-4 只是把它写得格外清楚:成绩越多,沉默也可以越理直气壮。
最醒目的成绩是模拟律师资格考试约进入前百分之十;同一报告把 GPT-3.5 写在大约后百分之十。这个对比很适合标题,也很容易诱使人误读。考试成绩测量的是模型在一组规定任务中的表现,不是职业资格,更不是对法律责任的理解。GPT-4 依然会编造事实,仍可能在复杂约束里犯下难以预料的错误。它的进步并没有把“不可靠”变成“可靠”,只是让错误更晚出现,也更容易夹在一段流畅、像样、甚至专业的回答里。这些数字来自 OpenAI 组织的评测设置,应被当作厂商自报的性能轮廓,而不是脱离提示、评分与脚手架仍成立的通用证书。
产品开放同样分了层次。ChatGPT Plus 订阅者可以在对话产品中使用 GPT-4;API 则以候补名单方式开放,并提供不同上下文长度的档位(发布材料中的 8K 与 32K)。价格按 token 计费,显著高于 GPT-3.5 系列。图像输入带来了另一种变化:使用者可以把图片和文字放进同一次交互,不必先自行把图表、文档或场景转写成一串文字标签。但首批图像能力仅限合作方测试,并未立刻向所有 API 用户完整开放;生产世界首先接到的仍主要是文本。多模态在 2023 年春天更像一张已经亮出的路线图。
安全与系统卡材料记录了另一条时间线。模型在拒绝有害请求、降低某些不当输出方面有过训练前后的对比;红队与缓解措施被写进发布配套文档。这些说明的是 OpenAI 当时的测试与约束努力,并不能被读成“已对齐完成”或“不会再出错”。外部研究者可以重复测试 API 输出,却无法从报告中核对模型规模、数据构成或训练算力。评测结果和内部机制自此成了两套不对称的证据。
GPT-4 留下的,是一张由发布方组织的性能轮廓:专业考试分数更高,图像输入已经演示,幻觉和推理错误仍在;安全训练改变了部分行为,却没有提供可靠性保证。ChatGPT 里的 GPT-4 与 API 里的 GPT-4 共享家族名称,但排队、速率、价格与是否立刻支持图像并不相同;合作方演示的视觉能力,不能直接写成“所有开发者当天可用”。对行业来说,它巩固了一种新的发布惯例——用榜单与产品档位说话,用沉默保护训练细节。3 月 14 日留下的第一张快照里,“更强、更贵、更不透明”同时成立。
影子可以量影长,却量不出机器里还藏着什么。
On 14 March 2023, OpenAI released GPT-4. The technical report is long. Readers can find a simulated bar exam, Olympiad problems, professional knowledge tests, and comparisons of the model before and after safety training. What lingers, though, may be what is missing: parameter count, training compute, dataset composition, architecture. Important machine-learning papers once opened the machine for peers. This time the machine stood behind evaluation tables and cast only a runtime shadow on the page.
That reading experience marks where GPT-4 already sat. It remained a research object, yet no longer belonged only to a paper community. ChatGPT had carried language models onto ordinary desks; capability had become product competition and tightly managed trade secrecy. OpenAI still spoke in the form of a technical report, but the report's job had changed. It no longer fully answered “what did we build.” It argued “what it can do,” explained “how we tested risk,” and drew a line around what would not be published.
The technical report, product notice, and system card formed three public texts on 14 March. Benchmarks were OpenAI's evaluation narrative; Plus and API waitlists were the access narrative; risks and mitigations were the safety narrative. They cite one another without substituting for one another. Outsiders can resample API behavior; they cannot resample the training mixture. That asymmetry became ordinary for frontier releases; GPT-4 simply made it unmistakable—more scores, more licensed silence.
The most quoted result placed GPT-4 around the top 10% on a simulated bar exam, with GPT-3.5 nearer the bottom 10% under the same reporting. The contrast travels easily through headlines and misleads just as easily. Exam scores measure performance on a defined set of tasks, not professional licensure and not legal responsibility. GPT-4 still fabricated facts and still failed inside complex constraints. Progress did not convert unreliability into reliability; it often delayed errors and wrapped them in fluent, professional-sounding prose. The figures are vendor-organized evaluations—performance contours drawn by the publisher, not universal certificates independent of prompts, grading, and scaffolding.
Access was staged as carefully as the tables. ChatGPT Plus subscribers could use GPT-4 in the chat product. The API opened through a waitlist, with context-length tiers described in launch materials (including 8K and 32K) and token prices well above the GPT-3.5 family. Image input was the other headline change: pictures and text could enter the same interaction without first being reduced to hand-written labels. At first, however, vision was limited to partner testing and was not fully open to every API customer. Production traffic mostly received text. Multimodality in spring 2023 looked more like a route map than a universal feature checklist.
Safety materials and a system card recorded another timeline: refusal behavior, reductions in certain harmful outputs, red-teaming, mitigations. Those documents describe OpenAI's then-current testing and constraints; they do not read as “alignment finished” or “errors ended.” Outsiders could retest API outputs; they could not audit scale, data mixture, or training compute from the report. Evaluations and internals became asymmetric evidence.
What GPT-4 left was a publisher-organized contour: higher professional-exam scores, demonstrated image input, remaining hallucination and reasoning failures; safety training that changed some behavior without guaranteeing reliability. For users, the practical ceiling rose on hard writing, coding, and analysis. For the industry, a release habit hardened—speak through benchmarks and product tiers, protect training detail with silence. Shadow length can be measured. The machine inside the shadow cannot.
展开完整事件档案人物、主题、模型与产品
- 人物
- —
- 模型
- gpt-4
- 产品
- chatgpt