OpenAI 发布 o3 与 o4-mini

首批能主动使用全套工具的推理模型

OpenAI 发布 o3 与 o4-mini,首次让推理模型在 ChatGPT 内自主调用搜索、代码执行、图像等全套工具。它把「推理」与「行动」首次系统性地结合到旗舰模型里。

时间2025 年 4 月 16 日 级别A · 行业级 组织OpenAI 状态已核验 · 1 个来源
嵌套推理环构成的紧凑立方体插画
o3 与 o4-mini 把强化学习推理进一步推向工具使用与自主搜索。 AI Chronicle

2025 年 4 月 16 日,OpenAI 发布了 o3 与 o4-mini。从名字上看,它们是 o 系列推理模型的新一代;但这次发布真正的变化,藏在「工具使用」这个词里。o3 和 o4-mini 是首批能在推理过程中主动调用工具的模型——搜索网页、执行代码、查看图像,把获取到的结果纳入自己的思考链,再给出最终答案。

在此之前,推理模型和工具使用是两条分开的路。o1 擅长「想」,但它的工具能力受限,需要用户自己把外部信息喂给它;工具型模型擅长「做」,但缺乏深度的推理规划。o3/o4-mini 把这两者合并:模型在内部推理时,可以实时决定「我现在需要查一下资料」,然后真的去搜索、去看、去算,再基于这些新信息继续推理。这是「边想边做」第一次在旗舰模型上成为完整闭环。

对开发者来说,这个变化是根本性的。过去构建一个智能体应用,需要自己编排模型调用、工具调用、结果回填的流程;现在,一次 API 调用就能让模型自己完成「思考+搜索+执行+回答」的全链路。开发者不用再写大量胶水代码,模型本身就成了一个能独立干活的执行体。这直接推动了 2025 年智能体应用的爆发。

o3 的能力提升也相当可观。在推理与编码基准上,它再次刷新纪录,把测试时计算的收益推到新高度;o4-mini 则用更低的成本提供接近旗舰的推理能力,让小开发者也能用上强推理模型。OpenAI 同时把这些能力深度集成进 ChatGPT,让普通用户也能体验到「模型自己上网查资料再回答」。

o3/o4-mini 发布的更深层影响,是定义了下一代模型的基本形态。此后的旗舰模型——无论是 GPT-5、Claude 4 还是各家竞品——都把「推理+工具+行动」的融合作为默认要求。「问答引擎」的时代过去了,模型开始成为「能独立完成任务的主体」。这种转变,比任何单次基准突破都更具结构性。

回看 2025 年 4 月,o3/o4-mini 的意义在于「融合」。它把思考与行动第一次完整地装进旗舰模型,让「模型自己干完整件事」从愿景变成日常。后来的 Agent 原生模型、自主智能体产品,几乎都站在 o3/o4-mini 划定的这条线上。一个模型不只是回答你,而是自己想办法、自己动手——这个画面在 2025 年春天之后,成为 AI 世界的常态。

On April 16, 2025 OpenAI released o3 and o4-mini. By name they were the new generation of the o-series reasoning models; but the real change was hidden in the phrase "tool use." o3 and o4-mini were the first reasoning models that actively call tools during reasoning—searching the web, executing code, viewing images, feeding the results into their thinking chain, then delivering a final answer.

Before this, reasoning models and tool use were separate paths. o1 was good at "thinking" but limited in tool ability, requiring the user to feed it external information; tool-capable models were good at "doing" but lacked deep reasoning and planning. o3/o4-mini merged the two: while reasoning internally, the model can decide in real time "I need to look this up," actually search, look, compute, and continue reasoning on the new information. "Think and act simultaneously" became a complete loop on a flagship for the first time.

For developers, the change was fundamental. Building an agent application used to require orchestrating model calls, tool calls, and result feeding yourself; now a single API call lets the model complete the whole "think-search-execute-answer" chain on its own. Developers no longer need to write piles of glue code; the model itself became an executor that works independently. This directly fueled the agent-application explosion of 2025.

o3's capability gains were substantial too. It broke records again on reasoning and coding benchmarks, pushing test-time compute to new heights; o4-mini delivered near-flagship reasoning at lower cost, letting small developers afford strong reasoning models. OpenAI also integrated these deeply into ChatGPT, letting ordinary users experience "the model searches the web itself and answers."

The deeper impact of the o3/o4-mini release was defining the basic form of the next-generation model. Subsequent flagships—GPT-5, Claude 4, and rivals—all made "reasoning plus tools plus action" a default requirement. The "question engine" era passed; models became "subjects that complete tasks independently." That shift was more structural than any single benchmark breakthrough.

Looking back at April 2025, the meaning of o3/o4-mini is "fusion." It packed thinking and acting completely into a flagship, turning "the model does a whole task by itself" from vision into routine. Later agent-native models and autonomous-agent products nearly all stand on the line o3/o4-mini drew. A model that does not just answer you but figures out how and does the work itself—after spring 2025, that picture became the norm in the AI world.

展开完整事件档案人物、主题、模型与产品
人物
模型
产品
来源

原始资料

  1. 01Introducing OpenAI o3 and o4-miniOpenAI · official

试试搜索