#23ON THE AI 30
Why AlphaFold Didn't Solve Protein Folding — Pushmeet Kohli, Google DeepMind & Sal Candido, Biohub thumbnail

LATENT SPACE · LATEST VIDEO

Why AlphaFold Didn't Solve Protein Folding — Pushmeet Kohli, Google DeepMind & Sal Candido, Biohub — Transcript

WATCH ON YOUTUBE 32m 5,676 words

Full Timestamped Transcript

00:00:00

So when people sort of say the protein folding problem has been solved like at at a conceptual level yes there might be sort of yes there has been there have been some advances we we are we don't know what is the actual true ground state that proteins take and what is the actual distribution of structure that proteins take. What we are trying to do is basically someone got a structure, deposited it in the PDB and we are trying to replicate and get the same structure. That's what we did, right?

00:00:35

And it just so happens that it's useful. But that does not mean that we have understood all of protein dynamics. >> Yeah. So, um, great to be here. Uh what an exciting morning. So many cool announcements. I think the future of bio science is being like announced right here. Um so this is the modeling session. Uh we are we are all modelers.

00:01:05

So of course it's natural for us to talk about data. Um so to start out with I think one of the things that I like to think about when it comes to data is uh or is like is there a scaling like how do you scale data properly? And uh this has brought me to this question about the bitter lesson but recast in the frame of data. So uh the bitter lesson for those of you who are not AI people is the statement that methods that scale win eventually. If you can just scale enough it wins. And so my question I'll start with Sal here is uh is there a better lesson for data?

00:01:45

Yeah, I mean for sure I think that um you you obviously need the the right data, right? Like in order for it to work. So I think like one misconception of scaling laws is that uh you know scaling laws are everywhere and they always exist. I think a lot of the work is actually finding that scaling law, right? Like so a lot of what we really do is trying to figure out uh what's a situation where if you put more compute into it, if you put more data into it, you actually get, you know, a better result out. And the reason to that's a great situation is then, you know, once that happens, you can kind of just turn the crank, right? Like becomes an engineering problem, which is something I like. But uh you know that it has to do with architecture, but it also very much has to do with data. So if you don't have the data with the right information statistics to solve the problem that you want, which is, you know, I think probably pretty obvious, like if you think about it for a second, like you're not going to get a model that has the capabilities and the understanding of you want. So you can

00:02:46

only really pull the information from the data that you have and then use that to generate like uh to generalize beyond, right? >> Yeah. So uh when it comes to data collection, how do you think about what types of data modalities are best? And uh maybe a kind of related question is do you think that modelers have a tendency to work on problems which are uh the data is available rather than maybe the data that or the problem that solves your goal most like maybe is has the most impact for translational medicine.

00:03:17

>> Oh for sure. So I think that like uh I won't speak for all modelers but like I'm lazy so I'm going to work with what's available and uh you know I think there is actually like a good side and a bad side to this and like the good side is that um you know I think like when you're doing like machine learning or like conventional machine learning you're really looking for like the most pristine highquality like data examples that you can find. But when you actually, if you're like me, you're actually someone who goes and you're kind of like rooting around in the back room looking for through people's like junk like what's there. So like a concrete example of this is when we train like when we train a protein language model. Uh we train on metagenomic sequences which are not the highest quality data. In fact, much of that data I can guarantee you isn't even like a real whole protein. And yet that makes the performance of the model go up for designing real proteins that work that make for understanding proteins

00:04:19

that like we know are things. And so you know that can so that's like a the positive side of it but that can lead you to like a negative place where you say like well let's just scale up the data that we can generate easily. And um I think that's not the necessarily the way to do it. I think that's one of the things that's really exciting to me about like what we're talking about here today with the BBI. Um, you know, it's really going out and saying like what is the data that we need to solve the problem. It's the the right resources but also the right community. Like I think that you know one thing that we do at uh or we try and do at Biohub anyway is to work in the open, work with the community and move the whole community forward. And I think that's a really critical piece for this because if you just have like people who are building models, you know, we're going to uh we're going to lean towards the data that exists. If you just have people who

00:05:20

generate the data, they're going to lean towards the things that can be generated. But like if you can work together as a community every step of the way in an open fashion, then you can actually figure out like what is the data that we need and then you can find that scaling law. Cool. Uh what do you think sit lesson for data? >> Yeah. So I I think the so I was at deep mind when Rich was uh with us and working uh and thinking about this this idea of uh the bitter lesson when Rich and I took my the the thing that I took from Rich's original sort of uh uh lecture that he actually gave at deep mind was not about the actual uh notion of whether data is useful or how should we think about machine learning. I think my uh sort of take was that he was talking about something more conceptual and the conceptual thing that he was talking about is sometimes um when we

00:06:22

are looking at problems we think about solutions in a very religious way. I'm a modeler or I'm a data generation person. And I think that is the bitter bitter lesson that if you approach the problem with that mindset, you might not succeed. The problem comes first and you should be flexible in terms of your solution space. You should try both things. You should see what uh you should try to understand what is the problem that you're trying to solve and if it requires modeling effort do the modeling effort and if it requires the data uh collecting more data then collect data. But sort of the thing that was that happened in machine learning in those in those times is people said we we are machine learning researchers. The data set is there. There here's some training data.

00:07:15

Here's some test data and we will just optimize the model and that is broken. Right? In the in the sense that if you if your eventual sort of goal is to solve the problem then you have to be able to look at both aspects of what goes into the process and what goes into the process is not just data or modeling it's also about sort of expertise. So the example uh that uh that I have is like with with AlphaFold um we looked at what was possible with current existing uh data sets because we did not have the the core expertise of now sort of going to um or even the resources by the way to uh sort of say let's augment the PDB uh by a significant sort of order of magnitude.

00:08:09

it was just not I mean the the amount of sort of investment that organizations and and scientists across the world had put in to construct that data was invaluable. So there you have to sort of uh focus on well the get getting the biggest bang out of your buck for uh by investing it in modeling. But in other areas when if you're sort of approaching say cell genomics then we took the same approach and said well what can you do with cell by gene if you're looking at sort of cell genomics and there it was very clear that we after a bunch of work that the data was not there yet to be able to build to go after that grand ambition of building the virtual cell.

00:08:53

And so it just brings people together to focus on what are the actual chi uh challenges if you want to advance science rather than being religiously following like uh advances on data generation or in modeling. >> Cool. Uh that's I really like that answer. So what do you think like the actionable takeaway is from this? Like I mean always uh define the problem first uh and then you know figure out what solution space to search over but more broadly uh like how should the community be thinking about this as we go forward for the next you know generation of uh basically trying to solve translational medicine.

00:09:33

>> Yeah I think the the the the advice that I give to anyone who is working who is starting in in the area is that think of yourself as a multid-disiplinary person. Understand the problem first. Why are you working on this problem? What are we trying to achieve? And and then think about what what is it that will be needed? Whether it's modeling, whether it will be compute scaling, whether it will be sort of data generation. So understanding of that problem sort of is extremely important and of course then you have to build sort of your expertise in the house. But um and and there are also it's not as if like you don't have constraints. Maybe there are constraints like you there is only a specific amount of data that you can sort of generate or there is um a specific model size that you can afford to sort of train.

00:10:22

Understand those constraints and and stop and try to fail fast and sort of look at what are the approaches that are going to be feasible in the long term in getting you to that intercept the the the intercept level the impact level that you're going after. >> Yeah. Okay. Cool. Uh so this actually leads me right into my next question which was um when I look at sort of the evolution of modeling uh let's say Alphaold 2 was uh essentially a work of art like a bunch of very carefully handcrafted features uh there was a lot of thought very intentional thought done to every part of the solution uh and then uh I I think you know some some of Google or Alphabet's uh work has been you know kind of still stays in that space but seems like the general consensus is to move to more scalable, more general strategies. Um, do you think that there's still a lot of like artisal craft solutions or problems that need that or is our resources better

00:11:22

spent on generally speaking um, you know, fixed resources, fixed money can go to compute, it can go to talent, it can go to data, you know, should we be focusing on scale first? So my uh again sort of let's uh approach it from first principles and think about when even so like when you think about like the arts and crafts of how do you construct a model to uh be better. I think it it was not an accidental thing right we did a lot of experimentation but there was a vision behind it that uh all this in scientific intuition that came from bioysics and biochemistry that those interactions that uh sort of amino acid sort of residues are not just doing their own thing they are being influenced by other residues. So let's bake that in like why like if you can uh if you have learned something from the scientific community use it use that information and try to sort of encourage the model and sort of give it that

00:12:23

unfair advantage that uh that it it has um and it it does basically makes the model much more data efficient because it doesn't have to replicate everything. I also have to sort of mention that um curating good data is an art in itself right it's not as if like you will say well put more data if you if you replicate the same amount of data you're not going anywhere so it's not just about big data it's about good data and actually understanding sort of the coverage of what data is necessary for making progress in the problem is in fact I would say a much interesting and much more challenging problem in itself Yeah. Cool. Well, do you uh looks like you have a thought. So, what do you think?

00:13:10

>> Yeah, I think um I mean I I very much agree. I think as you it really depends on you know as you're saying the the problem to be solved. So if you have kind of a smaller amount of data then certainly you know having more inductive bias in the model is going to help you right like and you actually need that to get the results on smaller scales of data. as you, you know, as you get more and more data, you know, you see that like sometimes the model can find things that like you didn't necessarily know about and sometimes that inductive bias if it wasn't exactly correct can hold you back, right? Like so there is definitely that like tipping point as things get better. But you know I guess like also I kind of object to your question too cuz I think there actually is a lot of uh I think there's a lot of craft actually to the scaling part of things as well. So you know you need very uh there's a lot of algorithmic work that actually goes

00:14:11

into taking these models and training them on more data putting more computing and making them bigger right like and we see that like not only is that from like an infrastructure perspective and you know making the models like run inference faster but you know we're actually seeing that like things are going beyond transformers to more bespoke architectures that can uh well I mean I don't think we're in a post-transformer world in any way, shape or form, but you know, we're modifying those architectures, right? Like in order to make them more fit for purpose and actually work better even at like internet scale data. So, you know, I think it's uh that's kind of the the project, right?

00:14:53

Like as more and more data comes in, it's not only, you know, helping to curate that data and select what is the next like batch of information that we actually need. you know, it's more information than data. How do you get that next batch of information that the models need? But then also like at every every step of the way, at every scale, what is the right architectures then to to get the most out of that data? >> Cool. So, uh if you think about the the state of protein, uh structure prediction right now, you know, the news will say protein structure prediction has been solved. Uh but then like if you talk to my friends they will all say we have so much like left to do here. Uh so many things like function dynamics uh and then design are still like just wide openen problems. Uh it's been fiveish years since Alphaold 2 was announced. Um progress has been made. Uh but I'm

00:15:53

curious uh push me what do you think the the like is there another big leap being made here? Is there block are there blockers in solving these problems? Um do we not have the right data yet? Do we not have the right algorithms um or ideas or is this like something that we can see evolving in the future? Uh and it's just kind of a matter of time. >> Yeah, I think science how it sort of operates is basically by isolating something and then making progress step by step. So when people sort of say the protein folding problem has been solved like at at a conceptual level yes there might be sort of yes there has been there have been some some advances but um I think it's to it's also like when you think about uh that narrative of proteins building being the building blocks uh proteins are not blocks and they're not sort of they don't act as as blocks right I I say proteins are the building blocks all time but actually

00:16:54

like I I don't believe in it right proteins are extremely complex uh they are disordered uh they their shape might change depending on the context and like what alphafold did and in fact like John John and I John jumper and I basically used to sort of discuss this like what are we trying to solve we we are we don't know what is the actual true ground state that proteins take and what is the actual distribution of structure that proteins take. What we are trying to do is basically someone got a structure, deposited it in the PDB and we are trying to replicate and get the same structure. That's what we did right and it just so happens that it's useful but that does not mean that we have understood all of protein dynamics. So I think the like at the at the at the top level it's easier to communicate that we have made

00:17:57

progress but the scientists among us and in this crowd the we know right and I think that's that's the really important thing for us to understand and have that common ground that a lot of work needs to be done and we have made progress and there's a lot to celebrate but uh let's not stop the funding of protein structure prediction and protein dynamics because we are just getting started. So what do you think the the biggest blocker if there's one thing you could wave a magic wand and say like we have more of blank that would accelerate one of those things function dynamics or design what would the magic wand be what would you wave into existence >> oh I I I have basically when when so I I'm a machine learning I my background is like quite eclectic I started as a security researcher then went into computer vision basian theory um and finally got into discriminative machine learning and deep learning and so on and u and then AI for coding and then

00:18:59

finally science. So when you asked me that question like my computer vision the computer vision researcher in me was super excited about crym microraphs. I was like what is this sort of PTB data I should be working at the source right I should be looking at the cryomm microraphs I I don't want those structures they must be sort of missing out all the data it's uh I should I should just directly operate at cryo microraphs now getting models that can scale at that level with the right amount of data and can extract all the dynamics and distributional information that is captured there would be an amazing thing we I tried I did but uh at it it requires more work right >> still work on it but you you believe that this is a route this this can give you >> yeah I think at some point of time maybe like people better than me would take a stab at it and uh and we'll get somewhere >> what do you think so >> yeah I think that um it's really interesting because I think these models

00:20:00

are um I mean they're they're quite useful but they're not exactly the problem that most people want to solve like they serve a very specific purpose and then you can also use them to do other like you can use them to design new proteins right like for example which is not the thing that you would kind of start with there but like uh so they're very useful but uh I kind of think of like the models that we build right now as uh you know you've got someone who's trying to like understand how like a bicycle works and you're like modeling like a spoke on it right like and you kind of have to like you know Those models can get better and better over time, but like a lot of what you really need to do is you need to move from models of like spokes to wheels to maybe whole bicycles because like that's actually what people want to understand.

00:20:49

Actually, you know, I think to continue the analogy, it seems like people want to use the model of the bicycle to design this part for a pickup truck or something like that, right? But um you it's I think that like as we like put these models more into like the particular context which they're in, right? Like then we'll actually be able to uh learn more about these interactions on a broader sense, right? And that's really I think where uh where things are going to be going. But I I agree with you. I think we should keep trying to work on folding models because I think they're going to keep getting better and better.

00:21:29

So, uh, with regards to design, um, there's this famous Fleman quote, uh, that which I cannot create, I do not understand. Um, now we're in a world where it's really easy to create things without understanding them at all. Uh so I I'm curious what your feelings your thoughts are about uh what like how important it is to actually have models which can help humans understand things versus magical black boxes which can effectively you know one shot you know a pico moler binder or something like that. I mean I think that like I mean we were designing things with magical black boxes long before like AI came around, right? Like so in some sense I think you know that's still useful but like I think that the understanding is like really important and I think that's actually one of the big things I think about a lot with AI because like you know there's so much for these models to learn and you know as like intelligence

00:22:31

is getting like cheaper and cheaper and more available you can deploy it to like learn more and more things about like what's going on in the world but then how do you actually you know pull that knowledge out of the machine so to speak right like and so it's something that like I can understand maybe that's just like my esoteric curiosity like I think there's so many things to learn >> I think lots of scientists really want to understand things and the end points are maybe not as important but then you know we are here to solve translational medicine as a as a problem right >> or I I mean I think it's like I think you do have to like the more you dig into things the more that you can like find the right way to to um to keep pushing them forward. So I think like one thing that's like really salient to me about this is I think the models actually have a lot more information in them than we know like you know so uh we've worked a lot on like interpretability for example for our models and you know you find a lot of

00:23:31

information in there which is uh you know I think people know that like protein language models like you know learn some notion of structure within their representations but we find information about uh you know some functions we find information about uh motions right like and so I think there's a lot in there still to be unlocked even from the models that we have now right like and I think that's important for us to understand how do these models work as we move up the you know as we raise the capability like at the end of the day like you expect like a world model to come out of like trying to compress all this information into the model, right? Like how does the model do its job? How does it design a protein? Well, you know, it's compressed all the information from evolution into this like model, right? Like so in addition to it, you know, being able to uh spit out something which is useful to us, there's certainly something to learn

00:24:31

just by looking at what's inside. >> Yeah. What do you think push me? Design versus understand. So, so I have a I have a I have a sort of u different take on this in the sense that I think when when we think about models I I think some level of understanding is necessary. So I'll say what level of understanding is necessary. So alpha fold actually is not perfect right it's not perfect it's 90 GDT alpha fold 2 rece was 90 GDT at that time on that sort of uh on on that set u but even if it was 95 GD but the PD the PLDDT score was completely unccalibrated who would trust it would just magically sort of give good answers but suddenly tell you here's very very confident and you will be working on it for the next one year and finding out it was

00:25:32

completely wrong. The calibration of the uncertainty measure was extremely important. So in that sense we do understand and we actually made a lot of progress in understanding how does alpha fold 2 behaves. Now that's different from how did it work and find the solution and there I agree with Sal that there's a notion of interpretability. Why did it work? Why did it give this answer? And I think if you think about it, AlphaFold 2 was at the highest level interpretable because if you uh if you look at how uh and this is the its ability to generalize and we did not sort of discover it before launching. When we launched Alpha Volt 2 and made the the weights available, people found out that it was a great disordered protein predictor. it could figure out like

00:26:34

which elements of the protein are disordered. So it just shows that it generalizes. Now interpretability asks the question interpretable by whom? If you're saying interpretable by a human rational system the with the cognitive and uh cognitive and computational limitations of the human brain then no alpha fold 2 is not interpretable. But if you are asking is alphafold too avail sort of interpretable to a much larger and more sophisticated model as to how it works maybe it is we just don't get it right but I think if as users of alphafold 2 we do need to understand what is it able to do and what it is not able to do so understanding the behavioral characterization and the strengths and limitations of the model and I think that is extremely important and we can't sort of just be using these models without having that behavioral characterization because otherwise it will sort of rather than help being

00:27:36

helpful harm us right >> so so your take is that interpretability strictly speaking isn't necessary but a way of ensuring the model is trustworthy so humans can make actionable decisions is the thing that people should be focusing on >> exactly and I also think that interpretability is also in the eyes of the beholder like who is interpreting it if it is a human scientist who is trying to interpret how the model is sort of going about and sort of how will it make this prediction. That's a different question from another much larger sort of LLM. If you if you give uh some of these large frontier models of the future access to the the activation layers and say okay tell me do you can you predict what alpha fold will do?

00:28:20

Maybe they will be able to predict and they will come up with a theory of how actually Alpha Fold 2 was interpreting and producing uh uh these results. >> Cool. Okay. Uh that's the two-minute warning. Um so I am one last question for both of you. Quick one. Um so there's a real chance that AI will make dramatic improvements in human health in the next you know you know immediate future. Um I like quantitative predictions. uh best guess how long until we start seeing uh AI results in the clinic um push me you want to start >> uh I mean the point is u AI results in the I think it's it's illposted it's it's not a u AI is being used today already in every part of the drug discovery process so if you if from that perspective it's already there right?

00:29:15

But if you are sort of asking me the question when will we see that 10x acceleration or 100x acceleration in the timelines then the other sort of question is what are we accelerating is it sort of uh doing lead optimization is it the target discovery is it uh sort of the pre-clinical uh or the talks sort of work so the uh I think over the next sort of few years and which This is why this effort that we announced today is extremely important. We need to tackle some of these hard challenges of biology and only then will we be able to get these true unlocks of acceleration that dramatically transforms how drug discovery will happen. So AI will be there all like in the clinic and will be impacting sort of things that go into the clinic all the time. But like the actual sort of larger

00:30:17

acceleration that will be unlocked only with a better understanding of the biolog biological models that this effort is trying to sort of create. >> Right. Thanks. So real quick question if we we're almost out of time but you >> Yeah. I mean, I'm tempted to like literally put on my bio hat to answer this question, but um you know, I think it's like it's really important uh to think about this, you know, as you were saying, like what does it mean to like push the field forward? Like the thing that we really want to see is those like outcomes and like so when will it be that like a drug that was made entirely by AI by an AI or something? I I I don't know. It's hard to it's hard to know that but like I do think we're going to see rapid progress here like very very quickly because like all these tools are being used and stuff like that and you know going back to your point about like the the basic you know we need basic research we need basic understanding you

00:31:18

know one thing I learned uh a long time ago in my career actually back at Google is like sometimes it t you know sometimes it's easier to approach a problem by looking at like what is it going to take to make a 10x improvement rather than a 10% improvement, right? And that's not because like that's necessarily an easier path. It's because it allows you to like take a broader view and see like some solutions that you know you haven't been approaching, right? Like and like what is actually the path to do it and you would go back and like look from first principles like how are we going to do this? How are we going to really push this? So I think like you know both like the uh both the approaches that like uh you know both approaches the the 10% and the 10x are valuable and we should do both. I'm just very happy that like at Biohub like we have a very uh I think a very beautiful and like lofty mission statement to to cure all disease, right? Like and if you

00:32:20

want to do that, you really need to go and figure out like what is the 10x approach and so I think it's uh it's a really great opportunity to be able to to go do that. >> Awesome. Thank you both. Thank you for being in the literal hot seat. Uh yeah, very interesting.

SHARE THIS TRANSCRIPT

Also on The AI 30

FULL CHART →

Transcript FAQ

Is there a transcript of "Why AlphaFold Didn't Solve Protein Folding — Pushmeet Kohli, Google DeepMind & Sal Candido, Biohub"?
Yes. This page has the full, timestamped transcript (5,676 words). Read it here or copy it with one click.
Why is this video on The AI 30?
It's the latest upload from Latent Space, which is #23 on The AI 30 with 291K subscribers.
Can I download this transcript as TXT or SRT?
Yes. Paste the video link into our free transcript tool to export it as plain text, SRT or VTT subtitles, or JSON.