IBM Watson 赢下《Jeopardy!》

大规模检索、排序与自然语言处理完成公开问答挑战

IBM Watson 在三场《Jeopardy!》比赛中击败两位冠军选手,展示机器处理双关、线索检索和答案置信度排序的能力。

时间2011 年 2 月 16 日 级别B · 领域级 组织IBM 状态已核验 · 2 个来源
纽约 Yorktown Heights 展出的 IBM Watson《危险边缘》比赛台
IBM 研究中心复原的《危险边缘》比赛台,中间屏幕代表 Watson。系统在 2011 年的电视赛中击败两位冠军选手。 Atomic Taco, CC BY-SA 2.0, via Wikimedia Commons

2011 年 2 月,Watson 坐上《Jeopardy!》舞台时,面对的是 Brad Rutter 与 Ken Jennings 两位顶尖冠军。节目给它的难题不只是“知道答案”。线索常有双关、文化暗示和省略表达;系统还要在几秒内判断自己有多确定,确定到什么程度才值得抢答,以及在 Daily Double 或 Final Jeopardy 中下注多少。答错不是一条无害的搜索结果,而会直接从分数里扣除。

舞台背后是一台占据十个机柜、包含 90 台服务器和 2,880 个处理器核心的计算机。竞赛版 Watson 不连接实时互联网,百科全书、词典、小说、戏剧与 Project Gutenberg 文本等资料预先装入系统。DeepQA 收到线索后,并行生成候选答案,再让数百种分析从类型、时间、地点、词语共现和文本证据等方向为候选打分。多个彼此独立的证据都指向同一答案时,置信度才会上升;整个过程约在三秒内完成。

这套设计最值得看的地方,是不确定性没有被藏起来。第一场 Final Jeopardy 的类别是“美国城市”,正确回答是 Chicago,Watson 却给出 “What is Toronto?????”。五个问号表示系统自己也不太相信这个答案,但下注机制仍让错误出现在数百万观众面前。机器不会因为电视直播而尴尬,工程团队却无法把边界留在实验室里:自然语言中的地理暗示,正好穿过了证据融合的缝隙。

三场比赛结束,Watson 以 77,147 美元胜出,Jennings 为 24,000 美元,Rutter 为 21,600 美元。这个分差证明 DeepQA 在《Jeopardy!》这套明确规则下确实达到了超越冠军选手的竞赛水平。与此同时,比赛也为能力划出了围栏:问题短、评分固定、语料预装、答案通常可由已有文本支撑。它不是不断变化的开放网络,更不是需要追责、解释和长期跟踪的医疗或企业决策环境。

Watson 品牌后来被带进医疗与企业市场,公众也因此容易用后续项目的成败重写 2011 年。两者应当分开。电视赛证明的是一种具体管线:先检索多个假设,让异质证据竞争,再根据置信度决定是否开口。这种思路后来仍能在检索增强问答中看到;奖杯本身却不能兑换成任何行业里的可靠性合同。

Watson 那三天最完整的记录,不是只留下冠军比分,也不是只留下 Toronto 的笑话。两者必须放在同一画面里:一边是机器在有界任务上的强大速度与证据融合,另一边是它公开承认却仍然说出口的低置信错误。真正可用的问答系统,既要会找到答案,也要知道什么时候不该按下抢答器。

When Watson took the Jeopardy! stage in February 2011, it faced Brad Rutter and Ken Jennings, two of the game’s greatest champions. The task was not merely to “know the answer.” Clues used puns, cultural references, and compressed wording. The system had to estimate its confidence within seconds, decide whether that confidence justified buzzing, and sometimes choose a wager. A bad answer was not a harmless search result. It subtracted money from the score.

Behind the stage stood a room-size machine: ten racks, 90 servers, and 2,880 processor cores. The competition system did not use the live internet. Encyclopedias, dictionaries, novels, plays, Project Gutenberg books, and other sources had been loaded in advance. After receiving a clue, DeepQA generated candidate answers in parallel and subjected them to hundreds of analyses—answer type, time, place, textual co-occurrence, and passage support among them. Confidence rose when independent lines of evidence converged. The full argument took about three seconds.

The revealing feature of this design was that uncertainty remained visible. In a Final Jeopardy category called “U.S. Cities,” the correct response was Chicago. Watson answered, “What is Toronto?????” The five question marks signaled that the system itself had little confidence, but the wager still placed the failure before millions of viewers. A machine cannot feel television embarrassment; its engineers could no longer keep the boundary inside the laboratory. A geographic implication in ordinary language had slipped through the evidence pipeline.

Across the three matches, Watson finished with $77,147, ahead of Jennings at $24,000 and Rutter at $21,600. The margin established genuine championship-level performance under Jeopardy! rules. The same rules drew a fence around the result: clues were short, scoring was fixed, the corpus was prepared, and answers were usually supported by existing text. The contest was not a continuously changing web, still less a medical or enterprise setting with liability, explanation, and long-term follow-up.

The Watson name later moved into healthcare and business products, which makes it tempting to rewrite the 2011 event through later successes and disappointments. The evidence should remain separate. The television system demonstrated a concrete pipeline: retrieve competing hypotheses, make heterogeneous evidence argue, and use confidence to decide whether to speak. That pattern still appears in retrieval-augmented question answering. A trophy, however, cannot be exchanged for a reliability contract in every other domain.

The complete record of those three days requires both images at once: the decisive score and the row of question marks after Toronto. One shows the speed and strength of evidence fusion in a bounded task. The other shows a low-confidence error that the system recognized yet still voiced. A dependable question-answering machine must do more than find plausible answers. It must also know when not to press the buzzer.

展开完整事件档案人物、主题、模型与产品
人物
模型
产品
ibm-watson
来源

原始资料

  1. 01IBM Watson — The Jeopardy ChallengeIBM · official
  2. 02Building Watson: An Overview of the DeepQA ProjectIBM Research · paper

试试搜索