
なぜ AlphaFold はタンパク質フォールディングを「解決」しなかったのか — Pushmeet Kohli(Google DeepMind)× Sal Candido(Biohub)
Why AlphaFold Didn't Solve Protein Folding — Pushmeet Kohli, Google DeepMind & Sal Candido, Biohub
Why AlphaFold Didn't Solve Protein Folding — Pushmeet Kohli, Google DeepMind & Sal Candido, Biohub
Latent Space: The AI Engineer Podcast
要約
Google DeepMind の Pushmeet Kohli と Biohub の Sal Candido が、モデラーの視点から生物学における AI のスケーリングとデータの在り方を議論したパネルです。正しいスケーリング則を見つけること、問題を起点に手法を選ぶこと、AlphaFold が構造予測の全てを解決したわけではないことが語られました。さらに、解釈可能性より信頼性と不確実性の較正が重要だという見方や、10x の改善を目指す発想の意義にも話が及びました。
- ●Sal は、スケーリング則はどこにでもあるわけではなく、それを見つける作業こそが重要だと述べている。見つかれば、あとはエンジニアリングの問題になる。
- ●Pushmeet は Bitter Lesson を、モデリングかデータ生成かといった手法へのこだわりを捨て、問題を起点に柔軟に解決策を選ぶ教訓だと解釈している。
- ●AlphaFold 2 には生物物理・生化学の知見を組み込んだ設計があり、データ効率を高めた。一方、データが増えるほど誤った帰納バイアスが足かせになりうると Sal は指摘する。
- ●Pushmeet は、AlphaFold が行ったのは PDB に登録された構造の再現であり、タンパク質の動態や構造の分布を理解したことにはならないとして、構造予測や動態研究への継続的な投資を訴えた。
- ●Pushmeet は、pLDDT のような不確実性の較正を含め、モデルの挙動の特性を把握することが重要だと述べている。人間にとっての解釈可能性は必須ではないという立場である。
- ●Sal は、Biohub の「全ての疾患を治す」という使命には、10% の改善ではなく 10x の改善を考える発想が必要だと述べている。
章立て
データ版 Bitter Lesson とスケーリング則
Bitter Lesson を踏まえ、データについてスケーリング則を見つける作業と、適切なデータの重要性を Sal が語る。
手に入るデータと質の低いデータ
モデラーは入手可能なデータに流れがちだが、メタゲノム配列のような質の低いデータでもタンパク質言語モデルの性能は上がる。コミュニティでの協働の重要性にも触れる。
問題を起点に考える DeepMind の教訓
Pushmeet が Rich Sutton の講義を引き、手法に宗教的にこだわらず問題から考える姿勢を説く。AlphaFold と仮想細胞の事例で、データ不足の判断を紹介する。
手作り設計とスケール、アーキテクチャの進化
AlphaFold 2 の科学的直観に基づく設計と、データの質の重要性、帰納バイアスの得失を議論する。Sal は、今はポストTransformer の時代ではなく、アーキテクチャが進化している段階だと述べる。
AlphaFold は何を解決していないのか
タンパク質は静的なブロックではなく、無秩序性や文脈依存の変化を持つ。AlphaFold は登録構造の再現であり、動態の理解にはまだ遠いと Pushmeet は述べる。
クライオ電顕から個別タンパク質を超えるモデルへ
クライオ電顕のミクログラフを直接扱う構想と、スポークから車輪、自転車全体へというモデルの拡張の比喩が語られる。
理解と信頼性: ファインマンの原理と解釈可能性
作れても理解できない時代に、モデル内部の知識の取り出しと、不確実性の較正や挙動の把握の重要性を議論する。
臨床への道筋と 10x の発想
創薬での AI 活用は既に始まっており、10x〜100x の加速には生物学の難題への取り組みが必要だと述べる。Sal は 10x の発想の価値を強調する。
解説記事
Latent Space のパネルセッションで、Google DeepMind の Pushmeet Kohli と Biohub の Sal Candido が、生物学向け AI をどう作るべきかを議論した。テーマは、スケーリング則の見つけ方、AlphaFold が残した課題、モデルの理解と信頼性、そして創薬への道筋である。モデラーの立場から、データと問題設定をめぐる率直な意見が交わされた。
データ版の Bitter Lesson
冒頭で司会は、「スケールできる手法が最終的に勝つ」という Bitter Lesson をデータの文脈に置き換えて問いかけた。Sal は、スケーリング則はどこにでも自然に存在するものではなく、計算やデータを増やせば性能が上がる状況を見つける作業こそが仕事の大半だと述べている。見つかれば後は「ハンドルを回す」エンジニアリングの問題になるが、そのためにはモデルが解くべき問題に必要な情報が、データに含まれていなければならない。
Sal はまた、モデラーは入手しやすいデータに流れがちだと認める。一方で、質の低いデータが役立つ例も挙げた。タンパク質言語モデルの学習に使うメタゲノム配列は、完全なタンパク質ですらないものが多いが、実在する機能的なタンパク質の設計の性能を上げるという。ただし、作りやすいデータを増やすだけでは不十分で、必要なデータを問い直すことが大切だと述べた。そのためにはモデルを作る人とデータを作る人が、オープンに協働することが重要だとしている。
問題が先、手法は後
Pushmeet は DeepMind 時代に Rich Sutton の講義を聴いた経験を踏まえ、Bitter Lesson を別の角度から読み解いた。自分は「モデラー」「データ生成の人」と手法に宗教的にこだわると失敗しやすく、先に来るのは問題だというのが彼の理解である。データセットを固定して最適化だけをする機械学習研究のやり方は、科学の問題を解くうえでは不十分だと述べた。
具体例として、AlphaFold では PDB を桁違いに拡張する専門性も資源もなかったため、既存データからモデリングで最大限の成果を引き出す方針を採った。一方、細胞ゲノミクスでは検討の結果、仮想細胞という大きな野心に見合うデータがまだ揃っていないと判断したという。新しく参入する人への助言は、多分野的な視点で問題を理解し、制約を把握し、早く失敗して長期的に実現可能な道を探ることだった。
手作り設計とスケールのバランス
AlphaFold 2 の手作り感のある設計について、Pushmeet はそれが偶然ではなく、残基同士が互いに影響し合うという生物物理・生化学の知見を意図的に組み込んだものだと説明した。こうした帰納バイアスはモデルのデータ効率を高める。また、同じデータを複製しても意味はなく、重要なのは量ではなく質とカバレッジだとも述べた。
Sal もデータが少ないうちは帰納バイアスが効くと同意しつつ、データが増えると、バイアスが正確でなければモデルの足かせになりうると指摘した。さらにスケールの側にも大量のアルゴリズム的な工夫があるとして、「ポストTransformer の世界にいるとは思わない」が、Transformer を目的に合わせて改造する動きは続いていると述べている。
AlphaFold は何を「解決」したのか
「構造予測は解決済み」という報道に対し、Pushmeet は、タンパク質は「ブロック」ではなく、無秩序で、文脈によって形が変わる複雑な存在だと述べた。AlphaFold がやったのは、PDB に登録された構造を再現することだった。それが有用だったのは確かだが、真の基底状態や構造の分布、動態を理解したことにはならないという。成果は祝福すべきだが、構造予測や動態研究への資金提供を止めるべきではなく、「まだ始まったばかり」だと訴えた。
次のブレイクスルーとして Pushmeet は、コンピュータビジョン出身の立場から、PDB の構造ではなくクライオ電顕のミクログラフを直接扱う構想に触れた。動態や分布の情報が含まれるが、実現には多くの作業が必要だと述べている。Sal は、現在のモデルは自転車のスポークだけを見ているようなものだと例え、今後は車輪、さらに自転車全体へと、文脈を含むモデルに広げていく必要があると語った。
理解と信頼性
ファインマンの「作れないものは理解できない」を引いた問いに対し、Sal は、モデルの内部には私たちが思う以上の情報があると述べた。タンパク質言語モデルは構造だけでなく、機能や運動に関する情報も表現に持っている。進化の情報を圧縮した世界モデルのようなものだと見て、その中身から学ぶ価値があるという考えである。
Pushmeet は、必要なのは完全な解釈可能性よりも信頼性だという立場を示した。精度が高くても pLDDT のような信頼度が較正されていなければ、誤った答えに一年を費やすことになりかねない。AlphaFold 2 が公開後に、無秩序領域の優れた予測器でもあると分かったことは、汎化性能の表れだと述べた。解釈可能性は「誰にとってか」で変わり、人間には解釈できなくても、将来の大規模モデルが活性化層を見て AlphaFold の挙動を説明できるかもしれないという見方も示した。
臨床への時間軸と 10x の発想
最後に、AI がいつ臨床に結果をもたらすかと問われた。Pushmeet は、創薬の各段階で AI は既に使われていると述べた。そのうえで、10x〜100x の加速が得られるかは、何を加速するのか(リード最適化、ターゲット探索、前臨床など)次第であり、生物学の難題に取り組む必要があるとした。Sal は、完全に AI が作った薬がいつ出るかは分からないが、進展は速いと述べた。また、10% ではなく 10x の改善を考えると視野が広がり、第一原理から道筋を見直せるという。Biohub の「全ての疾患を治す」という使命は、まさにその発想を求めるものだと語った。
まとめ
この議論から読み取れるのは、生物学向け AI では、計算やデータを増やすだけでなく、問題を起点に手法とデータ生成を組み合わせる姿勢が重視されている点である。AlphaFold の成功は出発点にすぎず、動態や無秩序性、システム全体の理解は未解決だというのが登壇者の見方だった。日本の研究者やエンジニアにとっては、既存データに最適化するだけでなく、必要なデータとは何かを問い直す視点と、精度に加えて不確実性の較正による信頼性の確保が実践上の示唆となるだろう。
文字起こし(英語・自動生成)
So, great to be here. What an exciting morning. So many cool announcements. I think the future of bioscience is being announced right here. So this is the modeling session. We are all modelers, so of course it's natural for us to talk about data. So to start out with, I think one of the things that I like to think about when it comes to data is, or is like, is there a scaling, like how do you scale data properly? And this brought me to this question about the bitter lesson, but recasting the frame of data. So the bitter lesson for those of you who are not AI people is the statement that methods that scale win eventually. If you can just scale enough, it wins. And so my question, I'll start with Sal here, is there a better lesson for data? Yeah, I mean, for sure.
I think that you obviously need the right data, right, like in order for it to work. So I think like one misconception of scaling laws is that, you know, scaling laws are everywhere and they always exist. I think a lot of the work is actually finding that scaling law, right? So a lot of what we really do is trying to figure out what's the situation where if you put more compute into it, if you put more data into it, you actually get a better result out. And the reason that's a great situation is then once that happens, you can kind of just turn the crank, right? Like it becomes an engineering problem, which is something I like. But that has to do with architecture, but it also very much has to do with data. So if you don't have the data with the right information statistics to solve the problem that you want, which is, you know, I think probably pretty obvious, like if you think about it for a second, like you're not going to get a model that has the capabilities and the understanding of what you want. So you can only really pull the information from the data that you have
and then use that to generate, like, to generalize beyond, right? Yeah. So when it comes to data collection, how do you think about what types of data modalities are best? And maybe a kind of related question is, do you think that modelers have a tendency to work on problems which are the data is available rather than maybe the data that or the problem that solves your goal most? Like maybe has the most impact for translational medicine? Oh, for sure. So I think that like I would speak for all modelers, but like I'm lazy, so I'm going to work with what's available. and, you know, I think there is actually, like, a good side and a bad side to this. And, like, the good side is that, you know, I think, like, when you're doing, like, machine learning or, like, conventional machine learning, you're really looking for, like, the most pristine, high-quality, like, data examples that you can find. But when you actually, if you're like me, you're actually someone who goes and you're kind of, like, rooting around in the back room
looking through people's junk, like, what's there? So, like, a concrete example of this is when we train, like, when we train a protein language model, we train on metagenomic sequences, which are not the highest quality data. In fact, much of that data, I can guarantee you, isn't even, like, a real whole protein. And yet that makes the performance of the model go up for designing real proteins that work, that make for understanding proteins that, like, we know are things. And so, you know, that can, that's like the positive side of it. But that can lead you to like a negative place where you say like, well, let's just scale up the data that we can generate easily. And I think that's not the way to do it. I think that's one of the things that's really exciting to me about like what we're talking about here today with the VBI. you know it's really going out and saying like what is the data that we need to solve the problem
it's the the right resources but also the right community like I think that you know one thing that we do at uh we try and do at biohub anyway is to work in the open work with the community and move the whole community forward and I think that's a really critical piece for this because if you just have, like, people who are building models, you know, we're going to lean towards the data that exists. If you just have people who generate the data, they're going to lean towards the things that can be generated. But, like, if you can work together as a community every step of the way in an open fashion, then you can actually figure out, like, what is the data that we need? And then you can find that scaling law. Cool. What do you think, Sal? A better lesson for data? Yeah, so I think I was at DeepMind when Rich was with us and working and thinking about this idea of the better lesson when Rich and I took
the thing that I took from Rich's original sort of lecture that he actually gave at DeepMind was not about the actual notion of whether data is useful or how should we think about machine learning, I think my sort of take was that he was talking about something more conceptual. And the conceptual thing that he was talking about is sometimes when we are looking at problems, we think about solutions in a very religious way. I'm a modeler or I'm a data generation person. And I think that is the bitter lesson that if you approach the problem with that mindset, you might not succeed. The problem comes first. And you should be flexible in terms of your solution space. You should try both things. You should try to understand what is the problem that you're trying to solve. And if it requires modeling effort, do the modeling effort.
And if it requires the data, collecting more data, then collect data. but sort of the thing that happened in machine learning in those sciences people said we are machine learning researchers the data set is there here is some training data here is some test data and we will just optimize the model and that is broken right in the sense that if your eventual sort of goal is to solve the problem then you have to be able to look at both aspects of what goes into the process. And what goes into the process is not just data or modeling. It's also about expertise. So the example that I have is like with AlphaFold, we looked at what was possible with current existing data sets because we did not have
the core expertise of now sort of going to or even the resources by the way to sort of say let's augment the PDB by a significant sort of order of magnitude it was just not the amount of sort of investment that organizations and scientists across the world had put in to construct that data was invaluable so there you have to sort of focus on well getting the biggest bang out of your buck for by investing it in modeling but in other areas that if you're sort of approaching say cell genomics then we took the same approach and said well what can you do with cell by gene if you looking at a sort of cell genomics and there it was very clear that we after a bunch of work that the data was not there yet to be able to build to go after that grand ambition of building the virtual cell. And so it just brings people together to focus on what are the actual challenges if you
want to advance science rather than being religiously following advances on data generation or in modeling. Cool. That's, I really like that answer. So what do you think, like, the actionable takeaway is from this? Like, I mean, always define the problem first and then, you know, figure out what solution space to search over. But more broadly, like, how should the community be thinking about this as we go forward for the next, you know, generation of basically trying to solve translational medicine? Yeah, I think the advice that I give to anyone who is working, who is starting in the area is that think of yourself as a multidisciplinary person. Understand the problem first. Why are you working on this problem? What are we trying to achieve? And then think about what is it that will be needed, whether it's modeling, whether it will be compute scaling, whether it will be sort of data generation. So understanding of that problem sort of is extremely important.
And of course, then you have to build sort of your expertise in the house. But and there are also it's not as if like you don't have constraints. Maybe there are constraints like you there is only a specific amount of data that you can sort of generate or there is a specific bottle size that you can afford to sort of train. understand those constraints and stop and try to fail fast and sort of look at what are the approaches that are going to be feasible in the long term in getting you to that intercept the intercept level the impact level that you're going after okay cool so this actually leads me right into my next question which is when I look at sort of the evolution of modeling let's say AlphaFold 2 was essentially a work of art, like a bunch of very carefully handcrafted features. There was a lot of thought, very intentional thought done to every part of the solution. And then I think, you know, some of Google or Alphabet's work has been, you know, kind
of still stays in that space, but it seems like the general consensus is to move to more scalable, more general strategies. Do you think that there's still a lot of, like, artisanal craft solutions or problems that need that? Or is our resources better spent on, generally speaking, you know, fixed resources, fixed money can go to compute. It can go to talent. It can go to data. You know, should we be focusing on scale first? So my, again, sort of let's approach it from first principles. And think about, even so, like when you think about like the arts and crafts of how do you construct a model to be better, I think it was not an accidental thing, right? We did a lot of experimentation, but there was a vision behind it that all this scientific intuition that came from biophysics and biochemistry, that those interactions, that sort of amino acid sort of residues are not just doing their own thing.
They are being influenced by other residues. So let's bake that in. Like if you can, if you have learned something from the scientific community, use it. Use that information and try to sort of encourage the model and sort of give it that unfair advantage that it has. And it does basically make the model much more data efficient because it doesn't have to replicate everything. I also have to sort of mention that curating good data is an art in itself, right? It's not as if like you will say, well, put more data. If you replicate the same amount of data, you're not going anywhere. So it's not just about big data. It's about good data. And actually understanding sort of the coverage of what data is necessary for making progress in the problem is, in fact, I would say, a much interesting and much more challenging problem in itself. Yeah. Well, it looks like you have a thought. So what do you think?
Yeah, I think, I mean, I very much agree. I think it really depends on, as you're saying, the problem to be solved. So if you have kind of a smaller amount of data, then certainly having more inductive bias in the model is going to help you. And you actually need that to get the results on smaller scales of data. As you get more and more data, you know you see that like sometimes the model can find things that like you didn't necessarily know about and sometimes that inductive bias if it wasn't exactly correct can hold you back right like so there is definitely that like tipping point as things get better but you know i guess like also i kind of object to your question too because i think there actually is a lot of uh i think there's a lot of crap actually to the scaling part of things as well so you know you You need very – there's a lot of algorithmic work that actually goes into taking these models and training them on more data,
putting more computing in them, making them bigger, right? And we see that not only is that from an infrastructure perspective and making the models run inference faster, but we're actually seeing that things are going beyond transformers to more bespoke architectures that can – Well, I mean, I don't think we're in a post-transformer world in any way, shape, or form, but, you know, we're modifying those architectures, right, like in order to make them more fit for purpose and actually work better, even at, like, Internet-scale data. So, you know, I think that's kind of the project, right? Like as more and more data comes in, it's not only, you know, helping to curate that data and select what is the next, like, batch of information that we actually need. You know, it's more information than data. How do you get that next batch of information that the models need? But then also, like, at every step of the way, at every scale,
what is the right architectures to get the most out of that data? Cool. So if you think about the state of protein structure prediction right now, You know, the news will say protein structural prediction has been solved. But then, like, if you talk to my friends, they will all say, we have so much left to do here. So many things like function, dynamics, and then design are still, like, just wide open problems. It's been five-ish years since AlphaFold 2 was announced. Progress has been made. but I'm curious, Pushmi, what do you think the, like, is there another big leap being made here? Are there blockers in solving these problems? Do we not have the right data yet? Do we not have the right algorithms or ideas? Or is this like something that we can see evolving in the future and it's just kind of a matter of time? Yeah, I think science
how it sort of operates is basically by isolating something and then making progress step by step. So when people sort of say the protein folding problem has been solved, like at a conceptual level, yes, there might be sort of, yes, there have been some advances but I think it also like when you think about that narrative of proteins being the building blocks proteins are not blocks And they don act as blocks right I say proteins are the building blocks all the time, but actually, like, I don't believe in it, right? Proteins are extremely complex. They are disordered. Their shape might change depending on the context. and like what AlphaFold did and in fact like John and I, John Jumper and I basically used to sort of discuss this like what are we trying to solve we don't know what is
the actual true ground state that proteins take and what is the actual distribution of structure that proteins take. What we are trying to do is basically someone got a structure deposited it in the PDB and we are trying to replicate and get the same structure. That's what we did, right? And it just so happens that it's useful. But that does not mean that we have understood all of protein dynamics. So I think at the top level, it's easier to communicate that we have made progress, but the scientists among us and in this crowd, we know. And I think that's the really important thing for us to understand and have that common ground that a lot of work needs to be done. And we have made progress, and there's a lot to celebrate, but let's not stop the funding of protein structure prediction and protein dynamics because we are just getting started.
So what do you think the biggest blocker, if there's one thing you could wave a magic wand and say, like, we have more of blank, that would accelerate one of those things, function, dynamics, or design, what would the magic wand be? What would you wave into existence? Oh, I have basically, so I'm a machine learning, my background is, like, quite eclectic. I started as a security researcher, then went into computer vision, vision theory, and finally got into discriminative machine learning and deep learning and so on, and then AI for coding, and then finally science. So when you asked me that question, like my computer vision, the computer vision researcher in me was super excited about cryo-EM micrographs. I was like, what is this sort of PDB data? I should be working at the source, right? I should be looking at the cryo-EM micrographs. I don't want those structures. They must be sort of missing out all the data. I should just directly operate at cryo-EM micrographs.
Now, getting models that can scale at that level with the right amount of data and can extract all the dynamic and distributional information that is captured there would be an amazing thing. I tried it, but it requires more work, right? So, we're going to, but you believe that this is a route. Yeah, I think at some point of time, maybe people better than me would take a stab at it, and we'll get somewhere. What do you think, Sal? Yeah, I think that it's really interesting because I think these models are, I mean, they're quite useful, but they're not exactly the problem that most people want to solve. They solve a very specific purpose, and then you can also use them to do other, but you can use them to design new proteins, right? Like, for example, which is not the thing that you would kind of start with there. But, like, so they're very useful. But I kind of think of, like, the models that we build right now as, you know,
you've got someone who's trying to, like, understand how, like, a bicycle works, and you're, like, modeling, like, a spoke on it, right? Like, and you kind of have to, like, you know, those models can get better and better over time, But, like, a lot of what you really need to do is you need to move from models of, like, spokes to wheels to maybe whole bicycles because, like, that's actually what people want us to understand. Actually, you know, I think to continue the analogy, it seems like people want to use the model of the bicycle to design this part for a pickup truck or something like that, right? But I think that as we put these models more into the particular context which they're in, then we'll actually be able to learn more about these interactions on a broader sense. And that's really, I think, where things are going to be going. But I agree with you. I think we should keep trying to work on holding models, because I think they're going to keep getting better and better. So with regards to design, there's a famous Feynman quote,
that which I cannot create, I do not understand. Now we're in a world where it's really easy to create things without understanding them at all. So I'm curious what your feelings, your thoughts are about what, like how important it is to actually have models which can help humans understand things versus magical black boxes, which can effectively, you know, one-shot, you know, a PicoMolar binder or something like that. I mean, I think that, like, I mean, we were designing things with magical black boxes long before, like, AI came around, right? Like, so in some sense, I think, you know, that's still useful. But, like, I think that the understanding is, like, really important. and I think that's actually one of the big things I think about a lot with AI because there's so much for these models to learn and as intelligence is getting cheaper and cheaper and more available, you can deploy it to learn more and more things about what's going on in the world.
But then how do you actually pull that knowledge out of the machine, so to speak? And so it's something that I can understand. Maybe that's just my esoteric curiosity. like I think there's so many things to learn I think lots of scientists really want to understand things and the end points are maybe not as important but then you know we are here to solve translational medicine as a problem right? I mean I think it's like I think you do have to like the more you dig into things the more that you can like find the right way to keep pushing them forward so I think like one thing that's like really salient to me about this is I think the models actually have a lot more information in them than we know. Like, you know, so we've worked a lot on, like, interpretability, for example, for our models. And, you know, you find a lot of information in there, which is, you know, I think people know that, like, protein language models, like, you know, learn some notion of structure within their representations.
But we find information about, you know, functions. We find information about motions, right? Like, and so I think there's a lot in there still to be unlocked, even from the models that we have now, right? Like, and I think that's important for us to understand. How do these models work as we move up, you know, as we raise the capability? Like, at the end of the day, like, you expect, like, a world model to come out of, like, trying to compress all this information into the model, right? Like, how does the model do its job? How does it design a protein? Well, you know, it's compressed all the information from evolution into this, like, model, right? Like, so, in addition to it, you know, being able to spit out something which is useful to us, there's certainly something to learn just by looking at what's inside. Yeah. What do you think, Prushmeet? Design versus understand So I have a sort of different take on this in the sense that I think when we think about
models, I think some level of understanding is necessary. So I'll say what level of understanding is necessary. So AlphaPold actually is not perfect. It's not perfect. It's 90GGT. AlphaFold 2 was 90 GDT at that time on that set. But even if it was 95 GDT, but the PLDGT score was completely uncalibrated, who would trust it? It would just magically sort of give good answers, but suddenly tell you, here's the answer, very, very confident, and you will be working on it for the next one year and finding out it was completely wrong. the calibration of the uncertainty measure was extremely important. So in that sense, we do understand, and we actually made a lot of progress in understanding how does AlphaFold2 behave.
Now that's different from how did it work and find the solution. And there I agree with Sal that there's a notion of interpretability. Why did it work? Why did it give this answer? And I think if you think about it, AlphaFold 2 was at the highest level interpretable because if you look at how, and this is its ability to generalize, and we did not sort of discover it before launching. When we launched AlphaFold 2 and made the weights available, people found out that it was a great disordered protein predictor. It could figure out which elements of the protein are disordered. So it just shows that it generalizes. Now, interpretability asks the question, interpretable by whom? If you are saying interpretable by a human rational system with the cognitive and
computational limitations of the human brain, then no, AlphaFold 2 is not interpretable. But if you are asking, is AlphaFold 2 sort of interpretable to a much larger and more sophisticated model as to how it works, maybe it is. We just don't get it. Right? But I think as users of AlphaFold 2, we do need to understand what is it able to do and what it is not able to do. So understanding the behavioral characterization and the strengths and limitations of the model. And I think that is extremely important. And we can't sort of just be using these models without having that behavioral characterization because otherwise it will sort of, rather than being helpful, harm us. Right? So your take is that interpretability, strictly speaking, isn't necessary, but a way of ensuring the model is trustworthy so humans can make actionable decisions is the thing that people should be focusing on. Exactly. And I also think that interpretability is also in the eyes of the beholder.
Like, who is interpreting it? If it is a human scientist who is trying to interpret how the model is sort of going about and sort of how will it make this prediction, that's a different question from another much larger sort of NLM. If you give some of these large frontier models of the future access to the activation layers and say, okay, tell me, can you predict what AlphaFold will do? Maybe they will be able to predict, and they will come up with a theory of how actually AlphaFold 2 was interpreting and producing these results. Cool. Okay, that's the two-minute warning. So one last question for both of you, a quick one. So there's a real chance that AI will make dramatic improvements in human health in the next, you know, immediate future. I like quantitative predictions best guess how long until we start seeing AI results in the clinic Vishneet you want to start I mean
the point is AI results in the clinic it's imposed it's not a AI is being used today already in every part of the drug discovery process So from that perspective, it's already there, right? But if you are sort of asking me the question, when will we see that 10x acceleration or 100x acceleration in the timelines, then the other sort of question is, what are we accelerating? Is it sort of doing lead optimization? Is it the target discovery? Is it sort of the preclinical or the TOCS sort of work? So I think over the next sort of few years, and which is why this effort that we announced today is extremely important, we need to tackle some of these hard challenges of biology. And only then will we be able to get these true unlocks of acceleration that dramatically transforms how drug discovery will happen.
so AI will be there like in the clinic and will be impacting sort of things that go into the clinic all the time but like the actual sort of larger acceleration that will be unlocked only with a better understanding of the biological models that this effort is trying to sort of create so real quick question if we are almost out of time yeah I mean I'm tempted to like literally put on my bio hat to answer this question. But, you know, I think it's like, it's really important to think about this, you know, as you were saying, like, what does it mean to, like, push the field forward? Like, the thing that we really want to see is those, like, outcomes. And, like, so when will it be that, like, a drug that was made entirely by AI, by an AI or something? I don't know. It's hard to, it's hard to know that. But, like, I do think we're going to see rapid progress here, like, very, very quickly because, like, all these tools are being used and stuff like that.
And, you know, going back to your point about, like, the basic, you know, we need basic research. We need basic understanding. You know, one thing I learned a long time ago in my career, actually back at Google, is, like, sometimes it, you know, Sometimes it's easier to approach a problem by looking at, like, what is it going to take to make a 10x improvement rather than a 10% improvement, right? And that's not because, like, that's necessarily an easier path. It's because it allows you to, like, take a broader view and see, like, some solutions that, you know, you haven't been approaching, right? Like, and, like, what is actually the path to do it? And you would go back and, like, look from first principles. Like, how are we going to do this? How are we going to really push this? So I think, like, you know, both, like, the approaches that, like, you know, both approaches, the 10% and the 10X are valuable, and we should do both. I'm just very happy that, like, at Biohub, like, we have a very, I think, a very beautiful and, like, lofty mission statement to cure all disease, right?
And if you want to do that, you really need to go and figure out, like, what is the 10x approach? And so I think it's a really great opportunity to be able to go through that. Awesome. Thank you both. Thank you for being in the literal hot seat. Thank you.
番組の概要欄(原文)
From the Bitter Lesson of AI scaling to the unsolved mysteries of protein folding, Google DeepMind’s Pushmeet Kohli and Biohub’s Sal Candido are rethinking what it takes to build AI that truly understands biology. In this special panel moderated by Brandon Anderson, they explore why AlphaFold’s breakthrough was only the beginning, why scaling compute and data alone won’t solve biology, and how the next generation of AI models could transform our understanding of proteins, cells, and human disease. We go deep on the future of AI-driven biology: finding scaling laws in biological data, the tradeoffs between scientific intuition and general-purpose architectures, why protein structure prediction is far from solved, and what it would take to build predictive models of living systems. Pushmeet reflects on the lessons behind AlphaFold, the limits of human interpretability, and whether future frontier models could understand other AI systems better than we can. Sal explains why protein language models may already contain scientific knowledge we haven’t unlocked, how biological modeling must move beyond individual proteins, and why achieving Biohub’s mission to cure all disease requires thinking in terms of 10x breakthroughs rather than incremental improvements. We discuss: * The Bitter Lesson for biology: why scaling compute and data isn’t enough * Why finding the right scaling law matters more than blindly increasing model size * How low-quality metagenomic data can improve protein language models * Why AI researchers optimize for available data instead of the most important scientific problems * Lessons from DeepMind on balancing modeling, data generation, and scientific expertise * Why building a virtual cell requires fundamentally different datasets * AlphaFold’s handcrafted architecture and the role of scientific intuition * Why good data matters more than simply having more data * Inductive biases, scaling laws, and the future of specialized AI architectures * Why we aren’t in a post-Transformer world, but architectures are evolving * Why AlphaFold didn’t actually solve all of protein folding * Protein dynamics, disorder, and the limitations of static structure prediction * How cryo-EM micrographs could unlock richer biological representations * Moving from models of individual proteins to whole biological systems * Feynman’s famous principle and why AI can now create things we don’t understand * The hidden biological knowledge inside protein language models * Why trustworthiness and uncertainty calibration matter more than full interpretability * Whether frontier AI models could interpret other AI systems better than humans * When AI could deliver 10x–100x acceleration in drug discovery * Why curing all disease requires thinking about 10x breakthroughs instead of 10% improvements Pushmeet Kohli — Google DeepMind * X: https://x.com/pushmeet * LinkedIn: https://www.linkedin.com/in/pushmeet-kohli-4838994/ Sal Candido — Biohub * Biohub: https://biohub.org/team/salvatore-candido/ * X: https://x.com/salcandido * LinkedIn: https://www.linkedin.com/in/salcandido/ Brandon Anderson — Moderator * LinkedIn: https://www.linkedin.com/in/brandon--anderson Timestamps 00:00:00 Introduction: The Bitter Lesson for Biological Data 00:01:00 Finding Scaling Laws and the Right Data for Biology 00:04:54 DeepMind’s Bitter Lesson: Solving Problems vs. Scaling Models 00:06:50 AlphaFold, Data Limitations, and Building the Virtual Cell 00:09:52 Handcrafted Architectures vs. Scaling Compute 00:11:53 Good Data, Inductive Bias, and Model Design 00:13:39 Beyond Transformers: The Future of AI Architectures 00:14:38 Why AlphaFold Hasn’t Solved Protein Folding 00:17:39 Protein Dynamics, Design, and Cryo-EM 00:19:09 From Individual Proteins to Whole Biological Systems 00:20:43 Feynman’s Principle: Creating Without Understanding 00:22:18 The Hidden Knowledge Inside Protein Language Models 00:23:48 AlphaFold, Trustworthiness, and Interpretability 00:25:56 Could AI Understand Other AI Models Better Than Humans? 00:27:43 When Will AI Revolutionize Drug Discovery? 00:29:47 Why Curing All Disease Requires 10x Breakthroughs Transcript Introduction: Is There a Bitter Lesson for Data? Brandon Anderson [00:00:04]: Great to be here. What an exciting morning. So many cool announcements. I think the future of bioscience is being announced right here. This is the modeling session. We’re all modelers, so of course it’s natural for us to talk about data. One of the things I like to think about when it comes to data is how to scale it properly. This has brought me to the question of the bitter lesson, but recast in the frame of data. The bitter lesson, for those of you who are not AI people, is the statement that methods that scale win eventually. If you can scale enough, it wins. So my question, starting with Sal, is: Is there a bitter lesson for data? Sal Candido [00:01:00]: For sure. You obviously need the right data in order for it to work. One misconception of scaling laws is that scaling laws are everywhere and they always exist. A lot of the work is actually finding that scaling law. A lot of what we do is trying to figure out: What’s a situation where, if you put more compute into it, if you put more data into it, you get a better result out? That’s a great situation because, once that happens, you can turn the crank. It becomes an engineering problem, which is something I like. That has to do with architecture, but it also very much has to do with data. If you don’t have data with the right information and statistics to solve the problem you want, you’re not going to get a model that has the capabilities and understanding you want. You can only really pull information from the data you have and use that to generalize beyond it. Choosing the Right Data: Availability vs. Scientific Impact Brandon Anderson [00:02:09]: When it comes to data collection, how do you think about which modalities are best? A related question: Do modelers tend to work on problems where the data is available, rather than the problems that best serve their goals or have the greatest impact on translational medicine? Sal Candido [00:02:32]: For sure. I won’t speak for all modelers, but I’m lazy, so I’m going to work with what’s available. There’s a good side and a bad side to this. The good side is that, when you’re doing conventional machine learning, you’re often looking for the most pristine, high-quality data examples you can find. But if you’re like me, you’re rooting around in the back room, looking through people’s junk to see what’s there. A concrete example is training a protein language model on metagenomic sequences, which are not the highest-quality data. In fact, much of that data, I can guarantee you, isn’t even a real, whole protein. And yet it makes the performance of the model go up for designing real proteins that work and understanding proteins that we know exist. That’s the positive side. But it can lead you to a negative place where you say, “Let’s just scale up the data that we can generate easily.” I don’t think that’s necessarily the way to do it. That’s one of the things that’s exciting to me about what we’re talking about here today with the BBI. It’s going out and asking, “What is the data that we need to solve the problem?” It’s about the right resources, but also the right community. One thing we try to do at Biohub is work in the open, work with the community, and move the whole community forward. That’s critical because, if you just have people building models, we’re going to lean toward the data that exists. If you just have people generating data, they’re going to lean toward the things that can be generated. But if you can work together as a community, every step of the way, in an open fashion, you can figure out what data you actually need. Then you can find that scaling law. The Problem Comes First: Modeling, Data, and Expertise Brandon Anderson [00:04:50]: Pushmeet, what do you think about the bitter lesson for data? Pushmeet Kohli [00:04:54]: I was at DeepMind when Rich was with us and thinking about this idea of the bitter lesson. What I took from Rich’s original lecture at DeepMind wasn’t about the specific notion of whether data is useful or how we should think about machine learning. My take was that he was talking about something more conceptual. Sometimes when we’re looking at problems, we think about solutions in a very religious way: “I’m a modeler,” or “I’m a data generation person.” I think that is the bitter lesson. If you approach a problem with that mindset, you might not succeed. The problem comes first, and you should be flexible in your solution space. You should try both things. You should understand the problem you’re trying to solve. If it requires modeling effort, do the modeling effort. If it requires collecting more data, then collect data. What happened in machine learning at the time was people saying, “We’re machine learning researchers. The dataset is there. Here’s some training data, here’s some test data, and we’ll just optimize the model.” That is broken in the sense that, if your eventual goal is to solve the problem, you have to look at both aspects of what goes into the process. And what goes into the process is not just data or modeling. It’s also expertise. With AlphaFold, for example, we looked at what was possible with existing datasets because we didn’t have the core expertise, or even the resources, to say, “Let’s augment the PDB by a significant order of magnitude.” The investment that organizations and scientists across the world had put into constructing that data was invaluable. So there, you had to focus on getting the biggest bang for your buck by investing in modeling. But in other areas, say cell genomics, we took the same approach and asked, “What can you do with cell-by-gene data?” After a lot of work, it was very clear that the data was not there yet to pursue that grand ambition of building the virtual cell. This brings people together to focus on the actual challenges of advancing science rather than religiously following advances in data generation or modeling. Brandon Anderson [00:08:23]: I really like that answer. What’s the actionable takeaway? Always define the problem first and then figure out which solution space to search over. But more broadly, how should the community think about this as we move toward the next generation of translational medicine? Pushmeet Kohli [00:08:47]: The advice I give to anyone starting in this area is to think of yourself as a multidisciplinary person. Understand the problem first. Why are you working on it? What are you trying to achieve? Then think about what’s needed, whether that’s modeling, compute scaling, or data generation. Understanding the problem is extremely important, and of course you need to build your expertise. There are constraints, too. Maybe there’s only a certain amount of data you can generate, or a certain model size you can afford to train. Understand those constraints, try to fail fast, and look at which approaches will be feasible in the long term to get you to the level of impact you’re aiming for. Scientific Inductive Bias vs. Scaling: The Craft of Building Models Brandon Anderson [00:09:52]: This leads right into my next question. When I look at the evolution of modeling, AlphaFold 2 was essentially a work of art, with a lot of carefully handcrafted features. There was very intentional thought put into every part of the solution. Some of Google’s or Alphabet’s work still stays in that space, but the general consensus seems to be moving toward more scalable, general strategies. Do we still need artisanal, craft solutions for certain problems? Or are resources better spent focusing on scale first? With fixed resources and money, you can invest in compute, talent, or data. How should we think about that trade-off? Pushmeet Kohli [00:10:48]: Let’s approach it from first principles. The art and craft of constructing a better model wasn’t accidental. We did a lot of experimentation, but there was a vision behind it. Scientific intuition from biophysics and biochemistry told us that amino acid residues are not just doing their own thing. They’re influenced by other residues. So let’s bake that in. If you’ve learned something from the scientific community, use that information. Give the model that unfair advantage. That makes the model much more data-efficient because it doesn’t have to relearn everything. I also have to mention that curating good data is an art in itself. It’s not as if you can just say, “Put in more data.” If you replicate the same data, you’re not going anywhere. It’s not just about big data; it’s about good data and understanding the coverage of the data necessary to make progress. That’s a much more interesting and challenging problem in itself. Brandon Anderson [00:12:20]: It looks like you have a thought, Sal. What do you think? Sal Candido [00:12:24]: I very much agree. It depends on the problem to be solved. If you have a smaller amount of data, having more inductive bias in the model is going to help you. You actually need that to get results at smaller data scales. As you get more and more data, sometimes the model finds things you didn’t necessarily know about, and sometimes that inductive bias, if it wasn’t exactly correct, can hold you back. There’s a tipping point as things get better. But I also object to your question a little because there’s a lot of craft in the scaling part of things as well. There’s a lot of algorithmic work that goes into taking models, training them on more data, putting more compute into them, and making them bigger. And that isn’t only from an infrastructure perspective or about making the models run inference faster. We’re seeing things go beyond standard transformers to more bespoke architectures. I don’t think we’re in a post-transformer world in any way, shape, or form, but we’re modifying those architectures to make them more fit for purpose and work better, even with internet-scale data. As more data comes in, the challenge isn’t only curating that data or selecting the next batch of information the models need. It’s also asking, at every step and at every scale, what the right architecture is to get the most out of that information. Why AlphaFold Didn’t Solve All of Protein Biology Brandon Anderson [00:14:36]: If you think about the state of protein structure prediction right now, the news will say that protein structure prediction has been solved. But if you talk to my friends, they’ll say, “We have so much left to do here.” Function, dynamics, and design are still wide-open problems. It’s been about five years since AlphaFold 2 was announced, and progress has been made. Pushmeet, do you think there’s another big leap coming? Are there blockers to solving these problems? Do we not have the right data, algorithms, or ideas? Or is it just a matter of time? Pushmeet Kohli [00:15:28]: Science operates by isolating something and then making progress step by step. When people say the protein-folding problem has been solved, at a conceptual level there have been advances. But think about the narrative of proteins being the building blocks. Proteins aren’t blocks, and they don’t act as blocks. I say proteins are the building blocks all the time, but I don’t actually believe it. Proteins are extremely complex. They’re disordered. Their shape might change depending on the context. John Jumper and I used to discuss this: What are we trying to solve? We don’t know the actual true ground state that proteins take, or the actual distribution of structures that proteins take. What we were trying to do was replicate a structure that somebody had obtained and deposited in the PDB. That’s what we did. And it just so happens that it’s useful. But that doesn’t mean we’ve understood all of protein dynamics. At the top level, it’s easier to communicate that we’ve made progress, but the scientists in this crowd know how much remains to be done. There’s a lot to celebrate, but let’s not stop funding protein structure prediction and protein dynamics, because we’re just getting started. Cryo-EM, Molecular Dynamics, and the Missing Information Brandon Anderson [00:17:39]: What’s the biggest blocker? If you could wave a magic wand and say, “We have more of this,” and that would accelerate function, dynamics, or design, what would you bring into existence? Pushmeet Kohli [00:17:54]: My background is quite eclectic. I started as a security researcher, then went into computer vision, Bayesian theory, discriminative machine learning, deep learning, AI for coding, and finally science. So when you ask me that question, the computer vision researcher in me gets very excited about cryo-EM micrographs. I thought, “What is this PDB data? I should be working at the source. I should be looking at cryo-EM micrographs. I don’t want those structures. They must be missing out on all the data. I should just operate directly on cryo-EM micrographs.” Getting models that can scale at that level, with the right amount of data, and extract the dynamics and distributional information captured there would be amazing. I tried it, but it requires more work. Brandon Anderson [00:18:57]: There’s still work to do, but you believe that’s a route that could give you that information? Pushmeet Kohli [00:19:00]: Yeah. I think at some point maybe some people better than me will take a stab at it, and we’ll get somewhere. From Protein Structures to Larger Biological Systems Brandon Anderson [00:19:07]: What do you think, Sal? Sal Candido [00:19:09]: It’s interesting because these models are quite useful, but they’re not exactly the problem that most people want to solve. They solve a very specific purpose, and you can also use them to do other things. For example, you can use them to design new proteins, which isn’t necessarily what you would start with. I think of the models we’re building now as someone who’s trying to understand how a bicycle works, but is modeling a spoke on it. Those models can get better and better over time, but what you really need to do is move from models of spokes to wheels to whole bicycles, because that’s what people want to understand. To continue the analogy, it seems like people want to use a model of the bicycle to design a part for a pickup truck. As we put these models into the particular biological context in which they’re operating, we’ll be able to learn more about these interactions on a broader scale. That’s where I think things are going. But I agree that we should keep working on folding models, because they’re going to keep getting better. Protein Design vs. Scientific Understanding Brandon Anderson [00:20:43]: With regard to design, there’s a famous Feynman quote: “That which I cannot create, I do not understand.” Now we’re in a world where it’s really easy to create things without understanding them at all. How important is it to have models that help humans understand things, versus magical black boxes that can effectively one-shot a picomolar binder or something like that? Sal Candido [00:21:18]: We were designing things with magical black boxes long before AI came around. In some sense that’s still useful. But understanding is really important, and it’s one of the big things I think about with AI. There’s so much for these models to learn. As intelligence gets cheaper and more available, you can deploy it to learn more and more about what’s going on in the world. But how do you pull that knowledge out of the machine so that I can understand it? Maybe that’s just my esoteric curiosity. I think there are so many things to learn. Brandon Anderson [00:22:05]: A lot of scientists really want to understand things, and the endpoints may not be as important to them. But we’re here to solve translational medicine as a problem, right? Sal Candido [00:22:18]: I think the more you dig into things, the more you find the right way to keep pushing them forward. One thing that’s really salient to me is that these models have a lot more information in them than we know. We’ve worked a lot on interpretability for our models, for example, and you find a lot of information there. People know that protein language models learn some notion of structure within their representations, but we find information about functions and motions as well. There’s a lot still to be unlocked, even from the models we have now. That’s important for us to understand as we raise their capabilities. At the end of the day, you expect a world model to emerge from compressing all this information into a model. How does it do its job? How does it design a protein? It has compressed information from evolution into that model. In addition to being able to produce something useful to us, there’s certainly something to learn just by looking at what’s inside. AlphaFold Confidence, Calibration, and Interpretability Brandon Anderson [00:23:48]: What do you think, Pushmeet? Design versus understanding? Pushmeet Kohli [00:23:53]: I have a different take in the sense that some level of understanding is necessary. Let me explain what level I mean. AlphaFold isn’t perfect. AlphaFold 2 had a GDT score of about 90 on that set at the time. But even if it had a GDT score of 95, if its pLDDT score were completely uncalibrated, who would trust it? Imagine that it magically gave good answers but told you it was very confident, and then you worked on it for the next year only to find out it was completely wrong. Calibration of the uncertainty measure was extremely important. In that sense, we do understand AlphaFold 2, and we made a lot of progress in understanding how it behaves. That’s different from understanding how it worked internally to find the solution. There I agree with Sal that interpretability asks: Why did it work? Why did it give this answer? At the highest level, AlphaFold 2 was interpretable in terms of its behavior and ability to generalize, and we didn’t discover everything about that before launching. When we launched AlphaFold 2 and made the weights available, people found that it was a great disordered-protein predictor. It could figure out which elements of the protein are disordered. That shows it generalizes. But interpretability asks another question: Interpretable by whom? If you’re saying interpretable by a human rational system, with the cognitive and computational limitations of the human brain, then no, AlphaFold 2 is not interpretable. But if you’re asking whether AlphaFold 2 is interpretable to a much larger, more sophisticated model in terms of how it works, maybe it is. We just don’t get it. As users of AlphaFold 2, we do need to understand what it can and cannot do. Understanding its behavioral characteristics, strengths, and limitations is extremely important. We can’t just use these models without that characterization, because otherwise, rather than being helpful, they can harm us. Brandon Anderson [00:26:53]: So your take is that interpretability, strictly speaking, isn’t necessary, but ensuring that a model is trustworthy so humans can make actionable decisions is what people should focus on. Pushmeet Kohli [00:27:04]: Exactly. Interpretability is also in the eye of the beholder. Who is interpreting it? If it’s a human scientist trying to interpret how the model makes a prediction, that’s a different question from giving another, much larger LLM access to the activation layers and saying, “Can you predict what AlphaFold will do?” Maybe those frontier models of the future will be able to predict that and come up with a theory of how AlphaFold 2 was interpreting and producing its results. When Will AI Transform Drug Discovery and Clinical Medicine? Brandon Anderson [00:27:43]: We have a two-minute warning, so one last question for both of you. There’s a real chance that AI will make dramatic improvements in human health in the immediate future. I like quantitative predictions. Best guess: How long until we start seeing AI results in the clinic? Pushmeet, do you want to start? Pushmeet Kohli [00:28:08]: I think “AI results in the clinic” is an ill-posed question. AI is already being used today in every part of the drug discovery process. From that perspective, it’s already there. But if you’re asking when we’ll see a 10x or 100x acceleration in timelines, then the next question is: What are we accelerating? Is it lead optimization? Target discovery? Preclinical work or toxicology? Over the next few years, and this is why the effort announced today is extremely important, we need to tackle some of the hard challenges of biology. Only then will we get the true unlocks of acceleration that dramatically transform drug discovery. AI will continue to be used in things that go into the clinic all the time, but larger acceleration will only be unlocked with a better understanding of the biological models this effort is trying to create. Brandon Anderson [00:29:43]: All right, thanks. Sal, a quick answer, if you can. We’re almost out of time. Sal Candido [00:29:47]: I’m tempted to literally put on my biohacker hat to answer this question. It’s important to think about what it means to push the field forward. What we really want to see is outcomes being affected. When will a drug be made entirely by AI? I don’t know. That’s hard to predict. But I do think we’re going to see rapid progress very quickly, because all these tools are already being used. Going back to your point about basic research and basic understanding, one thing I learned a long time ago in my career, back at Google, is that sometimes it’s easier to approach a problem by asking what it would take to make a 10x improvement rather than a 10% improvement. That’s not because the 10x path is necessarily easier. It’s because it allows you to take a broader view and see solutions you haven’t been approaching. You go back to first principles and ask, “How are we going to do this? How are we going to really push this?” Both approaches, the 10% and the 10x, are valuable, and we should do both. I’m very happy that at Biohub we have a beautiful and lofty mission statement: to cure all disease. If you want to do that, you really need to figure out what the 10x approach is. It’s a great opportunity to be able to go do that. Brandon Anderson [00:31:46]: Awesome. Thank you both. Thank you for being in the literal hot seat. Very interesting. This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit www.latent.space/subscribe