DeepMind 成立

把深度学习、强化学习与通用智能目标放进同一家公司

Demis Hassabis、Shane Legg 与 Mustafa Suleyman 在伦敦创立 DeepMind,目标是构建通用学习系统,再用它解决现实问题。

时间2010 年 级别B · 领域级 组织Google DeepMind 状态已核验 · 3 个来源
伦敦国王十字 Google DeepMind 总部入口
伦敦国王十字的 Google DeepMind 总部。DeepMind 于 2010 年在伦敦创立,2014 年被 Google 收购。 Gciriani, CC BY-SA 4.0, via Wikimedia Commons

2010 年,Demis Hassabis、Shane Legg 与 Mustafa Suleyman 在伦敦创立 DeepMind。公司想研究的不是某一款搜索、广告或语音产品,而是更一般的学习系统。这个目标几乎没有天然的交付日期:什么时候才算“通用”,又该用什么证据证明?如果只把它写在使命宣言里,它可以永远正确,也可以永远无法检验。

DeepMind 早期选中的试验场,恰好把这种模糊压缩成了清楚的反馈。电子游戏有画面、有动作、有分数,也允许同一局面被反复重来。研究者不必先教机器“飞船”“球拍”或“敌人”分别是什么,只要求它从像素中找到能提高奖励的动作。世界被缩得很小,学习却第一次可以连续地被观察:这一轮策略是否比上一轮多活了几秒,是否拿到了更高分。

这也解释了为什么 DeepMind 从一开始就把机器学习、神经科学、工程、数学和模拟放在同一个组织里。它要建造的不只是一个算法,还包括能快速生成经验的环境、能长时间训练的基础设施,以及让研究问题跨越多年而不被单个产品版本打断的预算。Google 在 2014 年收购 DeepMind,给了这套安排更大的算力和更长的时间,也让一家伦敦实验室进入大型科技公司争夺前沿研究能力的中心。

2015 年发表在 Nature 的 DQN 论文,是这条路线早期最清楚的切片。系统直接读取 Atari 2600 的像素,并用同一套核心方法学习多款游戏;经验回放降低连续样本之间的相关性,目标网络让训练更稳定。它没有获得关于现实世界的常识,也没有为每款游戏手写完整策略。可复查的进展来自一个更克制的承诺:在规则明确、奖励可见的环境中,深度神经网络可以和强化学习结合,从感知一路学到动作。

后来 AlphaGo、AlphaZero 与 AlphaFold 让 DeepMind 看起来像一条从游戏直达科学的成功直线,但这种回望会抹掉真正困难的部分。围棋有胜负,蛋白质结构有实验数据;医疗决策、社会系统和开放世界并不会自动提供同样干净的奖励。不同项目共享的是把问题改写成可计算、可比较、可迭代的工程习惯,而不是一台已经掌握所有领域的万能机器。

DeepMind 的创立因此更像一次组织实验:能否把一个过于宽广的智能目标,拆成一连串足够窄、足够诚实的测试,并持续为失败保留时间。游戏机屏幕上的分数没有证明通用智能已经抵达,却让一个原本只能写在宣言里的愿望,第一次有了可以逐局追问的进度条。

When Demis Hassabis, Shane Legg, and Mustafa Suleyman founded DeepMind in London in 2010, they did not organize it around a search feature, an advertising product, or a speech interface. The stated horizon was a more general kind of learning system. That ambition came without a natural delivery date. What would count as “general,” and what evidence could settle the question? Left at the level of a mission statement, it could remain inspiring forever and testable never.

The company’s early experimental environments did the opposite. Video games reduced an unruly idea to screens, actions, and scores. A system could try again from comparable states, and improvement could be counted. Researchers did not have to begin by defining every spaceship, paddle, or enemy. They could ask a narrower question: from pixels alone, can a learner discover actions that produce more reward? The world was deliberately small, but progress was visible one game at a time.

That choice also helps explain DeepMind’s institutional design. Machine learning researchers worked alongside neuroscientists, engineers, mathematicians, and specialists in simulation and computing infrastructure. The object being built was not only an algorithm. It was a set of conditions in which experience could be generated, training could run for long periods, and research questions could survive beyond a single product cycle. Google’s acquisition of DeepMind in 2014 enlarged the available compute and time horizon while drawing the London laboratory into the center of big technology’s competition for frontier research.

The 2015 Nature paper on deep Q-networks was an unusually clear slice of that program. DQN took raw Atari 2600 pixels and used the same core method across many games. Experience replay reduced correlations between successive samples; a target network helped stabilize learning. The system did not acquire common sense about the physical world, and it did not receive a complete hand-written strategy for each title. Its claim was more disciplined: in environments with explicit rules and rewards, deep neural networks and reinforcement learning could be joined from perception through action.

AlphaGo, AlphaZero, and AlphaFold later made DeepMind’s history look like a smooth road from games to science. It was not. Go supplies wins and losses; protein-structure work has experimental records against which predictions can be checked. Medicine, institutions, and open-ended social problems do not automatically provide such clean feedback. The projects share a habit of turning difficult questions into measurable improvement loops, not proof that one universal machine had already mastered every domain.

DeepMind’s founding was therefore an organizational experiment as much as a technical one: could an almost unbounded goal be divided into tests narrow enough to fail honestly, and could a company keep funding that sequence long enough to learn from the failures? An Atari score never demonstrated general intelligence. It did something more modest and more useful—it gave an otherwise ungradable ambition a progress bar that could be questioned after every game.

展开完整事件档案人物、主题、模型与产品
人物
Demis HassabisShane LeggMustafa Suleyman
模型
产品
来源

原始资料

  1. 01About Google DeepMindGoogle DeepMind · official
  2. 02Google buys UK artificial intelligence start-up DeepMindBBC News · archive
  3. 03Human-level control through deep reinforcement learningNature · paper

试试搜索