AlphaFold 破解蛋白质结构预测
CASP14 上接近实验精度的蛋白质结构预测
AlphaFold 2 在 CASP14 盲测的全部目标上取得 92.4 的中位 GDT_TS,许多预测接近实验结构精度。该成绩针对单链蛋白质结构预测,并不等于解决蛋白质动力学、相互作用或所有复合物结构。
92.4。
2020 年 CASP14 盲测里,DeepMind 的 AlphaFold 2 在全部目标上拿到约 92.4 的中位 GDT_TS——许多预测达到可与实验结构相比的精度,困难子集与个别结构域仍有落差。GDT_TS 衡量预测骨架与实验结构在多个距离阈值下的吻合;中位 92.4 不是每个残基的保修单,而是整体分布跨过了以往方法常见的鸿沟。氨基酸序列相对容易读出;三维结构通常要靠 X 射线晶体学、冷冻电镜或核磁共振,耗时长、成本高。计算结构预测已努力数十年。CASP 在实验结构公开前把目标序列发给参赛队伍,因此这份中位数带着较强的外部效度。
系统结合多序列比对中的共进化信号、残基对表示与结构模块,直接预测原子坐标,并给出 pLDDT 等置信度,让使用者知道哪些区段更可信。同源丰富的家族往往置信度更高;孤儿蛋白与快速演化区段更难。生物学家逐渐学会先看置信度色阶,再读骨架,再决定是否为实验买单。
2021 年代码与模型权重以研究使用为导向公开;随后 AlphaFold 蛋白质结构数据库扩展到超过 2 亿条预测结构,覆盖大量缺乏实验结构的序列。工作流层面的变化是:结构信息从稀缺实验产出,变成可检索的先验——先下载模型,形成关于结构域、界面候选与突变位置的假设,再排晶体学或突变实验的优先级。代码公开使方法可被修改与蒸馏;数据库让不会训练模型的研究者也能受益。
边界必须与成绩写在同一段。AlphaFold 2 针对的是单链(及有限条件下的)静态结构预测,不等于模拟折叠路径,不等于蛋白质动力学,不等于自动给出所有复合物与配体结合模式,更不等于替代药效与安全性实验。低置信区域、内在无序区段、稀有折叠与实验条件依赖的构象,仍会出错或含糊。数据库条目是预测,不是 PDB 实验条目的等价物。
结构模块输出坐标与置信度,数据库输出覆盖率,实验室输出最终裁决。三层分工在 CASP 之后逐渐稳定。AlphaFold 2 加速的是前两层;第三层从未外包给网络。分数表会旧,工作流留下了:先颜色,再骨架,再决定是否买单。
92.4.
In the 2020 CASP14 blind assessment, DeepMind’s AlphaFold 2 posted a median GDT_TS of about 92.4 across all targets—many predictions approaching experimental accuracy, with hard subsets and individual domains still lagging. GDT_TS measures backbone agreement with experiment at multiple distance cutoffs. A median of 92.4 is not a warranty on every residue; it is a distribution that crossed a gulf common to prior methods. Amino-acid sequences are comparatively easy to read. Three-dimensional structures usually require X-ray crystallography, cryo-EM, or NMR—slow and costly. Computational structure prediction had tried for decades. CASP releases target sequences before experimental structures are public, so the median carries real external validity.
The system combines coevolutionary signal from multiple-sequence alignments, pair representations, and a structure module to predict atomic coordinates, and it reports confidence measures such as pLDDT so users can see which stretches to trust. Rich families tend to score higher confidence; orphan proteins and fast-evolving regions struggle. Biologists learned a reading order: confidence color scale first, backbone second, experimental budget third.
In 2021 code and weights were released for research use; the AlphaFold Protein Structure Database later expanded beyond 200 million predicted structures, covering many sequences that lacked experimental entries. The workflow change is the real product: structural information shifts from scarce experimental output toward a searchable prior—download a model, form hypotheses about domains, candidate interfaces, and mutation sites, then prioritize crystallography or mutational assays. Code release enabled modification and distillation; the database extended benefits to researchers who never train models.
Boundaries belong in the same paragraph as the scores. AlphaFold 2 targets static single-chain (and limited related) structure prediction. It does not simulate the folding path, capture dynamics, automatically solve every complex and ligand pose, or replace efficacy and safety experiments. Low-confidence regions, intrinsically disordered stretches, rare folds, and conformation dependent on experimental conditions still fail or stay vague. Database entries are predictions, not equivalents of PDB experimental deposits.
The structure module outputs coordinates and confidence; the database outputs coverage; the laboratory outputs the final ruling. That three-layer division stabilized after CASP. AlphaFold 2 accelerates the first two; the third was never outsourced to a network. Score tables age. The workflow remains: color first, backbone second, then decide whether to pay for the experiment.
展开完整事件档案人物、主题、模型与产品
- 人物
- —
- 模型
- alphafold
- 产品
- —