Rich Sutton and Khurram Javed: Why AI Models Stop Learning, and How to Start It Again

Rich Sutton and Khurram Javed: Why AI Models Stop Learning, and How to Start It Again

Rich Sutton and Khurram Javed: Why AI Models Stop Learning, and How to Start It Again

Podcasts/Training Data/Rich Sutton, Khurram Javed

Podcasts/Training Data/Rich Sutton, Khurram Javed

Rich Sutton and Khurram Javed: Why AI Models Stop Learning, and How to Start It Again

Stream now on

Rich Sutton, who helped pioneer reinforcement learning and wrote the seminal AI essay The Bitter Lesson, has now cofounded Oak Lab with his former student Khurram Javed. Their goal: to build agents that continuously learn from their own experience rather than from us. Rich and Khurram argue synthetic data is "a big mistake." Their "big world hypothesis" is that the world is massively more complex than any agent or simulator, so approximations have to be updated continuously. Rich says LLMs represent roughly a quarter of intelligence, and that catastrophic forgetting is "totally curable" with the ideas behind their continual backprop algorithm. Their target: a trillion-parameter mind that keeps learning, stays coherent, and runs on 20 watts.

Watch Now

Transcript

Chapters

    Intro

    Rich Sutton: People think I have a radical point of view sometimes. They start questions saying how what I’m thinking is so different from everyone else. But I don’t see it that way at all. I see it as I’m thinking the ordinary way. It’s just everyone else that’s thinking a bit weird. And I mean that. It’s just the recent times people are thinking weird. Before there was all this AI craziness, you wouldn’t have to say “continual learning” because it wouldn’t make any sense to talk about learning that wasn’t continual. All learning is continual. We always act and we learn. That’s just the normal way of thinking. I’m not weird. The field is weird. They feel they need to call it continual learning. It’s just learning.

    Main conversation

    Sonya Huang: We are honored to have the great Rich Sutton with us here today. Rich, you invented reinforcement learning. You wrote the seminal textbook. You had the key students in the field, folks like Dave Silver. You wrote the essay, “The Bitter Lesson,” that I believe is the bible of the field. And you have just been one of the greats in propelling the field forward. So thank you for taking the time to join us today. Rich is joined by Khurram Javed, his co-founder and former student from the University of Alberta. The two of you have set off to found Oak Lab. I’m very excited to talk to you about that today.

    So for today’s session, we’re going to start talking about the bitter lesson, the state of the world as we know it today, whether LLMs will get us there or not. And then we’re going to transition to start talking about your research agenda and your plan for Oak. Rich, maybe take us back. I was going to start with the bitter lesson, but I actually want to start earlier than that. Decades ago, you decided to dedicate your career to reinforcement learning, to deep reinforcement learning in particular, and you established the University of Alberta as a bastion of that back when I think the field was very much in its infancy. What gave you the conviction to do that?

    Rich Sutton: What else are you gonna do? We’re trying to figure out the mind, and learning is a central part of the mind, and having a goal is a central part of the mind. Central part of intelligence. Yeah, so I was just doubling down on what I was always thinking.

    Sonya Huang: Did people think you were crazy at the time?

    Rich Sutton: It was a winter. It was an AI winter.

    Sonya Huang: What year was this?

    Rich Sutton: It was in 2003.

    Sonya Huang: Okay.

    Rich Sutton: And it’s kind of crazy, actually, the truth, because I was really sick. I was dying. I was actually dying of cancer in 2003.

    Sonya Huang: Oh my gosh!

    Rich Sutton: But I wasn’t quite dead—I’d been trying for a number of years and I wasn’t dead. I was in another remission. And so I said, well, I’m not dying. I haven’t succeeded in dying. So I might as well—it’s going on long enough, I might as well just try to get another job. And so I went to Alberta and started teaching there. And then in the end, I didn’t die. It’s kind of amazing. It’s like that. I’m joking about it now, but it was quite serious. And it’s an even more poignant question: Why did I continue to work on this research stuff when I only had a few months? I would always be reminded of what I think it’s Benjamin Franklin is supposed to have said, that if you ever wonder why someone is doing something, it’s almost always one of two things. It’s either habit or vanity. Okay? So I think it was probably true. Maybe it was a habit to just keep doing what I always was doing, or maybe it was vanity. I don’t know. I think it was more like a habit because I was dying.

    Sonya Huang: Wow. Divine intervention.

    Rich Sutton: Yeah, it’s always been easy for me to be very determined. And I’m going to go even longer on this answer.

    Sonya Huang: Please do.

    Rich Sutton: People think I have a radical point of view sometimes. They start questions saying how what I’m thinking is so different from everyone else. But I don’t see it that way at all. I see it as I’m thinking the ordinary way, it’s just everyone else that’s thinking a bit weird. And I mean that. It’s just the recent times people are thinking weird. If you look back, what people thought about the mind for even just a decade, you’ll find the kind of thoughts that learning is important, you’ve got to have a goal, and perception is important. And we have a—we are low-level beings, we are generating actions and perceiving data at a fast speed, and yet we have to think at higher levels. And go back a few, before there was all this AI craziness, you wouldn’t have to say “continual learning” because it wouldn’t make any sense to talk about learning that wasn’t continual. All learning is continual. It’s not a special phase. We always act and we learn. And that’s just the normal way of thinking. I’m not weird. The field is weird. They feel they need to call it continual learning. It’s just learning.

    Sonya Huang: I’m not weird. Everybody else is. That’s a good, good motto to live by.

    Alfred Lin: We’re going to have to send out an X post about that. We’re very happy that you lived on. The field is happy that you lived on. And thank you for pushing the frontier of AI.

    Rich Sutton: I’m really happy.

    Alfred Lin: Thank you for pushing the frontier of AI. I’m sure really happy. And thank you for all of that push. And you’ve been able to educate a lot of students who push the frontier as well. How’d you pick them, over the last 20, 30 years?

    Rich Sutton: Oh, well, you are giving me opportunities to be humble. I like to be humble and point out how all these great decisions just happen. And that’s the way I feel about students. I don’t feel that I choose them very well. Sometimes I’m lucky, sometimes I’m unlucky. I don’t feel I’m particularly good at picking my students. I’m looking at Khurram now. I think sometimes you end up with really great ones. David Silver picked me.

    Khurram Javed: Yeah.

    Rich Sutton: How was it that I got you, Khurram?

    Khurram Javed: Yeah, that was also—so I finished my master’s, not with you, and I was planning to join the industry. And then we were collaborating on a project—which also just started organically. There was something I worked on that Rich was in a meeting, then they mentioned that I worked on it. So I got pulled into it. We started collaborating. It went really well. I felt so happy with that collaboration. Rich also felt really good about it. And then six months down the road, we had made some progress, and it just made sense to convert that into a thesis proposal. So at no point did I apply, at no point did I ask, should you be my PhD advisor? We worked together, then we decided this would be a pretty good thesis. And then after that, I applied for the PhD.

    Sonya Huang: Life works in unexpected ways. Take us to 2019. You wrote “The Bitter Lesson,” which has become the modern tome. 2019 was a funny time to be writing that piece because ImageNet was 2009, AlphaGo was 2015. What caused you in 2019 to reflect and to write that? Because it was before the current kind of scaling paradigm around large language models had taken off, but it was after deep learning had really proven itself.

    Rich Sutton: Well, it was a long time coming. As “The Bitter Lesson” expresses, it’s something that you can observe for a long time, for many decades. And it’s definitely at least as much due to the round of symbolic AI, which I lived through. It’s all about not getting distracted by trying to put in your human knowledge and just paying attention to what the problem needs and how you can scale with computation. I know I wrote versions of it at least a year before, and I gave talks. I gave a talk a year before, and it wasn’t a particular response to the moment. It was a particular response to my long experience, different people trying to think in different ways about how you can make smart systems.

    Sonya Huang: What is the essence of the bitter lesson?

    Rich Sutton: The essence of the bitter lesson.

    Sonya Huang: It may be the phrase that I hear used the most in my meetings these days. “Is this bitter lesson pilled? Is it not bitter lesson pilled?”

    Rich Sutton: Yeah.

    Sonya Huang: I would imagine, given the popularity of the phrase, it’s probably been tortured and misused in different ways that you didn’t originally intend it. So what do you think? What is the essence of it? And where do you think people go wrong in their attempt to understand it?

    Rich Sutton: Yeah, you’re making me think about X now. And I recently made a post where I tried to do the bitter lesson in 26 words.

    Sonya Huang: [laughs]

    Rich Sutton: It goes something like, “Don’t be distracted by human knowledge as AI traditionally has been many times. Instead, focus on learning methods that will scale with computation, like search and like learning. So it’s really all about focusing on algorithms and improvements. It’s not saying you don’t need fancy algorithms. You need fancy algorithms, but you want fancy algorithms that will scale with computation.”

    Alfred Lin: Rather than scaling with data.

    Rich Sutton: Rather than scaling with human input. Yeah. And then the question, if I can anticipate it, yeah, what about large language models?

    Alfred Lin: Are they consistent or inconsistent with your essay?

    Rich Sutton: And I’ve thought about this. And I think there’s another X post about it. But the conclusion is that it’s both a positive example and a negative example of the bitter lesson. First, large language models enabled enormous scaling with computation, and you could just drink in the internet and scale so much. So it was a way of getting a much more capable system just by methods at scale.

    Then after that, as you go on further, it eventually gets limited by that information. The internet is finite and it’s hard to get more examples. And the world is big, and the world is massively bigger than everything we’ve stored on the internet. And so in the end, it seems like it could be—I guess that would be a positive example of when we relied too much on human knowledge and it eventually holds us back.

    Sonya Huang: Hmm. Can I just push on this a little bit?

    Rich Sutton: Yeah.

    Sonya Huang: It seems like a lot of what the foundation model labs are working on right now is synthetic data generation in order to get us beyond the fossil fuel that is the existing human internet. Is synthetic data generation, as part of this LLM scaling paradigm, is that a general method that leverages computation?

    Rich Sutton: No.

    Sonya Huang: Why?

    Rich Sutton: That’s just a big mistake.

    Sonya Huang: Why?

    Rich Sutton: Well, it’s such a big—maybe it’s the next big lesson. It’s been floating around Alberta for five or 10 years. And we call it the big world perspective or the big world hypothesis. Khurram, who eventually wrote it up as a paper. There’s a little paper called “The Big World Hypothesis.”

    Khurram Javed: The big world is that the world is infinitely big. There are infinitely many things to learn, and you can have people generating these synthetic datasets, but there would always be more things to learn. And because of that, if you could just learn from experience, if you could remove the humans from the loop, then you would have systems that can do everything. Because the world is big, there are many tasks that we want them to do, and they would be able to do anything by learning from their own experience.

    Going back to the synthetic data question, too, who decides what’s good synthetic data and what’s bad synthetic data? Because I can write a program that can output a lot of synthetic data which would hurt programs. So right now I would say humans decide, and that’s the bottleneck where okay, you can have humans deciding how to generate these datasets, but you need human experts who know what’s a good dataset and what’s a bad dataset for that approach to scale. So it is bottlenecked by humans.

    Sonya Huang: Doesn’t my loss curve decide how much better did I get with this dataset versus that dataset?

    Khurram Javed: Right, but if all the engineers, OpenAI, Anthropic, or all the big neo labs, their engineers went on vacation, who would generate the synthetic data? That’s the question. It doesn’t come from agents’ experience. It’s not something that the agent is generating itself. Some human has to decide what is the right synthetic data to generate. And that requires human expertise.

    So for example, if you want a system to do something very challenging from a physics point of view, maybe you want a drone that flies with echolocation like a bat, for example, what’s the right synthetic data for that? I think you would need to hire domain experts to go figure out what is the right data and generate it, and then maybe you would be able to learn from that. But the domain expert has to exist first. So we are bottlenecked by human expertise at that point.

    Sonya Huang: But you can have infinite synthetic worlds. The existing world is finite.

    Khurram Javed: But let’s go back to the echolocation thing, right? That’s what I want. I want a drone that can localize itself and move with echolocation. If that’s my goal, that’s a robot generating its own experience, so it could totally learn from its own experience. But it wouldn’t be able to. It doesn’t matter how much synthetic data you generate. It doesn’t matter if you generate synthetic data that captures 50 different universes. It will not allow you to do that task without humans figuring it out first.

    Rich Sutton: But first, just say the synthetic data is wrong. I mean, it won’t be correct. It’ll be a synthetic world. It won’t be the real world. And it will matter. The world is incredibly complex. If you write a little program, because it’s going to be a little program that’ll generate the synthetic data, it’ll be a very—it’ll be a small world.

    Sonya Huang: Hmm.

    Rich Sutton: So for example, what’s important to me is what’s going on in your mind right now. Okay? And you’re saying, why don’t I get some synthetic data to tell me what’s going on in other people’s minds? No, there’s no way we can have synthetic data for other people’s minds, and other people’s minds matter to us. I talked to you guys about investing today, so I care what’s going on in your minds. But how can I get synthetic data on such a thing? Really, you can’t even get synthetic data on anything. You can’t get synthetic data on how the drone is going to interact with its environment in the physical world, and the friction and wear in the motors of this robot. The world is infinitely complex and any simulation of it is microscopic.

    The big world hypothesis, let’s say what it is, is that the world is massively more complex than your mind, than any agent. And this is obvious because the world contains many other agents. So because the world is massively complex, there’s no way you can do anything that might claim to be optimal or perfect. You’re going to be imperfect and you have to have approximations, and those approximations will be severe. And so because of that, that is ultimately the reason why we have to continue learning. If you want to think of it as a reason, we have to continue learning because we’ll encounter some particular part of this immense world and we’ll have to learn. And approximation that’s tuned to the part of the world we’re in, not to all the other parts that we’re not in.

    Sonya Huang: Yeah. I’m going to push on this one more time. And sorry, I’m being argumentative for the sake of being argumentative.

    Rich Sutton: No, good to be.

    Sonya Huang: I’m trying to understand. My understanding is that the newest cohort of self-driving car companies, many of them were primarily trained in sim, and then they have to do some sort of post-training, I guess, to make sure they work in the real world, but that it’s been a very effective pipeline.

    Khurram Javed: Yeah, so I think the important question to ask here is how many engineers were involved in building that simulation? And are we ready to say that the only problem worth solving are those where we can hire a large team of engineers to first make a simulation? And I’m sure they had to do multiple iterations where they made the simulation, they learned in it, they realized there was a sim-to-real gap that was not acceptable, then they fixed it. So there is this human in the loop fixing the simulation. Like, they’re getting feedback from the real world—humans—and then they’re fixing the simulation. Why can’t we just remove the human and let the agent do it itself?

    Rich Sutton: And then when it actually drives, again, something unexpected will happen.

    Sonya Huang: And that’s when they actually want to learn from experience. So your point is there’s just so much more data that’s going to come from experience than there possibly can be from humans curating, creating data.

    Khurram Javed: Yeah, and I think there is obviously value in learning from simulation. And there is a way of doing it. The agent can learn a model from its own experience, and when the agent learns this, it’s much better because if the model is incorrect, it can fix it by continual learning. If the humans are making the simulator, then the model only gets updated when the humans figure out that something’s wrong. So yes, planning is important. The agents should learn from simulators, but simulators they make themselves.

    Sonya Huang: Okay. I want to move to another part of the bitter lesson: removing human knowledge. From your essay, quote, “Seeking an improvement that makes a difference in the shorter term, researchers seek to leverage their human knowledge of the domain, but the only thing that matters in the long run is the leveraging of computation.”

    And if I—your former student Dave, with AlphaGo and AlphaZero—for me, that was an example of a triumph of removing human priors. Did that result surprise you? Why or why not?

    Rich Sutton: Of course it made me very happy. It made me feel vindicated. You know, it could have gone either way. It wasn’t that—because prior knowledge can help. There’s nothing wrong with prior knowledge. And I say this right at the very beginning of “The Bitter Lesson.” I say there’s no reason why there has to be a conflict between prior knowledge and then learning knowledge. You can put some prior in there and then start learning. There’s no reason in principle why these have to be opposed. In fact, they’re all about knowledge. Life is gaining knowledge and having knowledge. And why are these—somehow nature and nurture became enemies. But really, prior learning is what you already—and then you learn more. And they should be friends.

    But as I say at the beginning of “The Bitter Lesson,” in practice they have been enemies. In practice, people who had an affection for existing human knowledge ended up wanting that to win. And so they wanted to minimize or dismiss learning. And so now I’m sure your sense of me is that I’m someone who loves learning and wants to dismiss prior knowledge, but I’m really someone who’s interested in the mind. The mind is, you have prior knowledge and then you get more, and then once you’ve gotten more, then that becomes your prior knowledge as you get more and more and more.

    And these two things work together. I end up appearing to be someone who’s interested in learning primarily because all the rest of the world is talking about all you need is enough knowledge. You don’t need to learn. Large language models are oh, we’re gonna put all this knowledge into the system, and the large language model will not learn when it runs. It’s talking to people, it’s interacting. It is absolutely—the weights never change. So I am not the weird one. It’s you guys that are the weird one that think that that’s possible, that you could possibly claim they can make a PhD-level experience and expertise out of something that doesn’t learn at all anymore. So I’m not the weird one.

    Alfred Lin: [laughs] So your recommendation is let the algorithms run for a much, much longer period of time before feeding ...

    Rich Sutton: Continually learn.

    Alfred Lin: Before you feed it data, prior data. Or drip the prior data along the way, prior knowledge.

    Rich Sutton: So both are important, but in the long run, you’ve got to gain and structure the gaining of new knowledge. That’s what matters in the long run. And as you are doing this yeah, there’ll be some that you got previously. How will it work—you know, look into the future when we have intelligent robots. Will we have them all learn from scratch? Or will we copy them and ask them to keep learning from wherever they are? I mean, they’ll be digital and it’ll be easy to copy them. And so instead of having this huge thing where we’re spending zillions of dollars to retrain them from the internet, we’ll just copy the agent and keep learning from there.

    And so in some sense, the prior knowledge will be—should be dismissed, because you’re just going to copy it from the previous robot.

    Alfred Lin: So why don’t you just describe for us what you think a machine or a computer that learns from experience looks like?

    Rich Sutton: Well, it could look like a robot, but it also could live entirely on the internet. You could, for example, have routing of packets through the internet, and do that in a way that’s sensitive to experience and becomes better over time. Or you can interact, be the user interface that’s interacting with people, like on your phone or on your computer, and it becomes better over time. Yeah, like an intelligent assistant has to become better over time, has to know what you want.

    Sonya Huang: Would your contention be that the current paradigm of the popular LLM-based assistants, would your contention be that these are not experiential learners or continual learners? And if so, what is the fundamental gap?

    Rich Sutton: Are you serious?

    Sonya Huang: [laughs]

    Rich Sutton: I mean, obviously they have ...

    Sonya Huang: They learn memories about me. They’re doing some in-context learning.

    Rich Sutton: Their weights never change.

    Sonya Huang: And by the way, is a small number of the weights changing sufficient, or do you need all the weights to be changing?

    Rich Sutton: Well, so think of all the structuring and generation of new concepts that went into creating the large language models. All that is the weight learning. And you want to be able to continue doing that. You don’t want that to happen just once.

    Alfred Lin: Is another way of saying it is we do too much pre-training and post-training before we launch the models? They don’t learn after that.

    Khurram Javed: The only point with the big disagreement is we don’t let them learn after that.

    Alfred Lin: Yeah, we don’t let them learn after that.

    Khurram Javed: We can do as much pre-training as we want. That’s okay. Post-training is fine. But then when I’m using the model, it stops learning. You can give it more context. You can change the state of the model by giving it more context. And so it has already learned that if the state is different, if the state says something new, then it’ll use that to make the next prediction. But the model is unlearning.

    Sonya Huang: Cursor’s tab autocomplete model. It does get updated based on ...

    Alfred Lin: Those models, those weights change.

    Khurram Javed: Those weights change. Those are two examples, Cursor’s Tab, and I think the Composer, they were also updating. Those are two examples of continual learning. But it can be much better. So the way they do it, as far as I understand, is a lot of people are using Tab, they collect all this data, so coming from millions of users or thousands of users, and then they do one update of the policy from this batch data. So this could work, but imagine I want to teach this model something specific. I don’t want to fight with 100,000 other people about what they want to create and teach their models. I want to teach my model something very specific, and I want to do it to my version of the model. I don’t care about the shared knowledge that the model has coming from other people. And so it’s a very inefficient way of doing it.

    Sonya Huang: It seems like the way that this is currently done is that there’s fundamental skills maybe that are learned in the weights that are common to everybody. And then there’s personalization that happens in the form of context, right? Yeah. Is that not the right mental model for how learning should work? Should all the context live in the weights themselves?

    Khurram Javed: So context can be in the state, too. It could be both, but you still need to be able to update the weights. So if I give you an example, some really good use studies are with human disabilities. When humans go through something that changes their mind or some sensors, you can see them adapt. So for example, we have proprioception, we have internal sensors that tell us how the body is positioned, and we use this for walking. There are cases where people lose this ability completely, and then they can’t walk at all because that is literally the foundation of their walking policies. It is ingrained in the brain. But then over the course of two, three years, they can learn to walk again by looking at their feet. So visual feedback through that. So the brain is insanely plastic in the sense that it can learn a lot of things. Something that has been true for 20 years, when it stops being true, it can go and update that and get rid of that. And that is the capability I think that’s extremely useful that we would want in our systems.

    Sonya Huang: What is there for us to learn from how human babies or animals learn, and how much inspiration do you take from that?

    Rich Sutton: Well, we take a lot of inspiration. We don’t take it as a requirement that the AI has to behave like the natural system, like babies or people or animals. But it’s a source of inspiration but not constraint from animal learning.

    Sonya Huang: To be consistent with the bitter lesson? Where do you think we should most seek to draw inspiration from the way that biological learning works that is not present in today’s systems?

    Rich Sutton: I feel like I’m just giving opinions now, but they’re just obvious opinions. So I think it’s apparent that no animal learns by supervised learning, because we don’t get examples of how our muscles should twitch. And that’s our output.

    Sonya Huang: But all of school is supervised learning.

    Rich Sutton: No, absolutely not. But even if it was, school is a tiny fraction of what we learn. We learn to see and we learn to walk and we learn how the world works. But even in school, no one tells us how we should twitch our muscles.

    Sonya Huang: The knowledge skills I acquire were from supervised learning in school.

    Rich Sutton: So I don’t want to say that learning from others, transmission from others, is not important. It’s extremely important. And language is extremely important. But what are we missing? There is no supervised learning. There’s no targets that are given to us. You hear the right answer is, you know, what’s the capital of France? And we know the answer is Paris. Okay, but no one tells me how I should pronounce “Paris.” You say the answer is Paris and I listen to you and I hear your words, and I will make some other muscle motions to produce the answer “Paris.” It’s not literally supervised learning. Anyway, yeah, so I think it’s really true. I mean—well, anyway, the first thing is the school is irrelevant. Squirrels don’t go to school and learn that.

    Sonya Huang: They might!

    Rich Sutton: Animals don’t learn that way. And school is a very special thing that even we didn’t have up until, I don’t know, a few hundred years ago. But it’s not part of the essence of intelligence. And it’s a distraction to think of that as your primary example of learning, is this thing which we didn’t do as animals.

    Alfred Lin: I wish you had been around to tell my parents that. They forced me to go to school and deal with all the structure.

    Sonya Huang: The thing is, squirrels are wonderful at jumping off trees, but squirrels can’t prove math theorems. And if I want to learn how to prove a math theorem, I go to school.

    Rich Sutton: Yeah. They also don’t have DVDs and iPods. There are a lot of things—they can do things that we can’t do. But math theorems? Yeah, and they don’t play chess. Sort of like Moravec’s paradox. There are these advanced things that we think of as really intelligent, but they’re sort of easy for computers to do, as opposed to all these regular things that are hard, like moving and seeing with attention and everything. I think supervised learning is a good thing. I like to look for obvious things. No one tells us how to twitch our muscles by giving us examples, because they couldn’t possibly, because we have had to twitch our muscles. We’ve had to figure that out.

    Sonya Huang: Yeah.

    Khurram Javed: And their answer would be wrong, right? So, if I moved my mouth and my tongue and my vocal cords exactly the same way that Rich does to pronounce Paris, I’m sure a very different sound would come out. So in some sense, Rich or no one knows the right way of producing a sound with my body. Only I know that.

    Sonya Huang: Yeah. It seems to me that many of the most raw sensory motor capabilities, especially related to movement in the physical world, I agree with you that that seems something that is inherently learned from experience. It seems to me, though, that there are higher levels of abstraction that bring us closer to what makes humans great. And much of that doesn’t live in this low level of sensory motor learning. Does your world model span sensorimotor learning all the way up?

    Rich Sutton: Yeah, that’s the ambition. Absolutely. And squirrels, by the way, can do some enormously abstract things.

    Sonya Huang: What’s the coolest thing a squirrel can do?

    Rich Sutton: Well, it can always get into your bird feeder, no matter what obstacles you put in the way. It can find new ways to jump and climb and do lots of things.

    Alfred Lin: They calculate trajectories pretty well. Animals are pretty good at understanding the physical world without the mental calculations that we think we are doing when we think about launching ourselves into space.

    Khurram Javed: Breaking a fall, they can do it in real time the right way to prevent injuries.

    Sonya Huang: Okay, fair enough.

    Rich Sutton: I think it’s just a question of degree between—and I like to think that animals, other animals, are very close to humans. I think it’s hubristic to try to emphasize what we do differently, how we’re different from animals. It’s better to see the commonalities. And I think we are just a question of degree. It’s degree, and of course, society and culture give us big advantages. Language gives us big advantages.

    Alfred Lin: Can I just push on some of this?

    Rich Sutton: Yeah, good.

    Alfred Lin: Because I want to back up Sonya. So I believe animals and children learn from experience, and do incredible things learning from experience. And when my son was two or three or four, I’m like, “Wow, this is really interesting that my son can learn these things without nobody really teaching him how to do these things.” But at the same time, what Sonya is saying is like what makes human uniquely human, to be able to go to outer space, build a rocket, those are not things that are learned 100 percent from experience, because before you launch the rocket, you actually have to abstract thinking through it in a way that is not learned from quote-unquote “experience,” because you don’t know if it’s going to work or not. You have to imagine it. How do we teach a machine to imagine things that were not available before? That’s probably the thing that we’re trying to push on because we’re not quite understanding that.

    Rich Sutton: We’re absolutely going to agree with you there. You have to be able to plan. You have to be able to imagine.

    Alfred Lin: Yeah.

    Khurram Javed: Would you say that humans 1,000 years ago, before they had done all—most of the things that we were talking about, were they as intelligent? If, for example, someone from that era was exposed to this new culture, would they be able to get the same skills and start doing useful things?

    Alfred Lin: Even over the last 10,000 years, I don’t think the human brain has evolved that much.

    Rich Sutton: Fundamentally the same machine.

    Alfred Lin: Fundamentally the same machine, but we’ve built up 10,000 years of knowledge.

    Khurram Javed: Yes.

    Alfred Lin: And I get to learn 10,000 years of knowledge by going to school through supervised learning, right? And I get all that much, much faster than trying to learn through experience.

    Khurram Javed: So I think you’re totally right. So we would want our systems to learn from experience, and part of their experience would be getting exposed to our culture and then learning about our culture. They should learn from that. That’s all good. But let’s talk about when someone goes and does a paradigm-shifting thing. So everyone gives the example of Einstein, but I think there are many examples. Learning is that too, like looking at learning things versus programming things. So when these paradigm shifts happen, I would say it’s a human who has accumulated all this knowledge, and then from their experience, they’re building new abstractions, they’re planning with them, and then they’re discovering new knowledge. And that skill of coming up with new abstractions and then learning world models and planning with them, that skill is totally missing in our current systems. And you can expose this at the edge of human knowledge, but you can also study this problem at the sensory motor stream level.

    Rich Sutton: So we’re not arguing with the principle. We need to form abstractions so we can reason at a high level. You guys are coming close to doing that thing that I said we should never do, which is argue is prior knowledge important or gaining knowledge important? That’s what you guys just said. You said you’re still going to have to learn things. And you were saying, oh, I can get things from my culture and from prior knowledge. But these should not fight for each other.

    Alfred Lin: On the exact thing around paradigm shifts, how do we create a machine that understands when to shift the paradigm?

    Khurram Javed: I think through its experience, right? So it would have to, through its own experience. It can’t rely on human knowledge, because we’re assuming the humans see one paradigm and we want a different way of looking at things. And so through its experience it has to find something that is better. Maybe it generalizes better and makes better predictions. Maybe it’s better in some other ways, but it has to be through its own experience.

    Rich Sutton: The big challenge that we don’t see in our field, the ability we don’t see in our field yet, is the ability to learn a model and then plan with a model. We can do the math things, and we can do AlphaGo because the games, we know the model, we know how the moves work. And in math, we know what the operators are. We know lean will take us from one state of knowledge to the state of the proof to the next state. But if we have to learn the models, there are no—I’m going to say it, it’s probably maybe a weird example, but a counterexample, but I’m going to say there’s no instances of learning the model and then planning with the model in our field.

    Khurram Javed: At least not with self-discovered abstractions. So there are people who say, “I’m just going to learn a model of what happens in the next second or next millisecond,” but that’s not how our models work. Our models are more abstract. Our models are quite different.

    Sonya Huang: So one of the things I like about what you’re doing here is you’re not just sitting around pontificating or lamenting the state of the world as it is. You’re very action-oriented; that’s why you’ve started a company. So let’s start talking about that a bit. In 2022, Rich, you laid out a very specific 12-point plan, the Alberta Plan for AI Research. Maybe tell us about that.

    Rich Sutton: So the Alberta Plan came about because we did have general ideas, but we also needed to convert them into smaller chunks. And so the 12 steps are the attempt to crystallize particular chunks. There’s a very important early step, step two, which is continual deep learning. And we think that one is almost the most important because it unlocks everything else. If you could do continual deep learning, you could then continually update your model of the world. And then if you knew how to do the abstraction right—and the second half of the steps are all about how to get the abstractions right. And not only abstractions, right? What do I mean? I don’t mean get the right abstractions, because no one can say what the right abstractions are. That depends on the world that you’re in. Your agent would have to learn the correct abstractions for whatever world it’s in. And so maybe those are the two key things: You have to find the right abstractions, and then you have to be able to do continual deep learning.

    Khurram Javed: I think that a lot of the people in the field realize that we need models, we need to plan with them, but the abstractions tell us what the model should be conditioned on. So what should the model predict? What are you going to do, and then something is going to happen? And more importantly, where would that come from? So I really like the example of elite athletes. If you ask elite athletes about how they do certain things, they would have weird, niche terminologies for doing very specific things. You know, “I do this thing,” and they would have a name for it if they communicate. Sometimes they don’t even have a name for it if they’re just doing it alone. So how did they come up with those abstractions? That’s in some sense a crucial thing that’s missing that the later half of the Alberta Plan answers.

    Sonya Huang: Can we talk about the continual deep learning part? Is it an algorithmic gap that exists today, or is it just a practical deployment infrastructure data privacy gap? Because if I wanted to do—call it naive—updating of weights based on user interaction, I can do that today, right? And so what, in your opinion, is the biggest thing that we’re missing to get to continual deep learning?

    Khurram Javed: Yeah, so it’s absolutely an algorithmic gap. You can do the naive thing, but then you’ll see all sorts of problems. For example, if you say, “I’m gonna take one sample and then I’m gonna update my whole model with that one sample,” you will run into this problem that now all of the previous knowledge in the model, it’s impacted negatively. And the way currently we get around this is exactly what Cursor does. They don’t use one example, they use a large batch coming from a lot of users. So in use cases where you can have that, you can do continual learning. But in most use cases you don’t have that. Most use cases you have a single stream of data, and then if you apply it to the naive thing, it just completely destroys your prior knowledge in a very destructive way.

    Rich Sutton: Catastrophic forgetting.

    Sonya Huang: Yeah.

    Rich Sutton: But it’s totally curable. You have to have the right algorithm.

    Sonya Huang: What’s the cure?

    Alfred Lin: What is the cure?

    Rich Sutton: Well, first you need to do what we call step-size optimization. And it means every weight in your network has to have a separate step size. So some will move fast, some will move slow. And you will have to meta-learn these step sizes for each weight. Most of your network will have weights that have tiny step sizes, so then when you train on a new example, they don’t get destroyed.

    Sonya Huang: Mm-hmm.

    Rich Sutton: It happens just to the right places. And then secondly, you have to use some form of generate and test, which is in feature space. So you come up with new features or new units, and without following gradients—because gradients is a very slow process. You only move in a direction if you know it’s the helpful one. And that’s always going to be very slow, and doesn’t give you a path to grow more and more complex and to have sustained learning. You need to have something that just proposes a bunch of new units and then goes from there, I guess.

    So there is a specific thing I can say to make it at least concrete, which is to say we have this algorithm called continual backprop, which we published in Nature a couple years ago, and it is exactly like backprop, but you also plant new seeds of units that are newly initialized with random weights. Backprop only has random weights at the beginning of time, and then as you go on, all that randomness, all that variety from the randomness gets used up. And with continual backprop, you keep injecting a bit of randomness, a bit of generate and test, a bit of generate, and then the operation of backprop is the tester. So you need that, and if you put those together really well, I think you’ll have a new generation of massively superior continual deep learning. And that’s what we hope to do in the next couple years.

    Sonya Huang: Wonderful. Do you think that these algorithms can be applied to the current state of affairs with people scaling LLMs and trying to get them to do continual learning without catastrophic forgetting?

    Khurram Javed: Yeah, absolutely. I think it’s—so I don’t think that you could take an existing model and say, “I’m going to just start updating it with these algorithms,” because these algorithms meta-learn how to learn. So really you have to say, “I’m going to learn from scratch.” So let’s say learn a new foundation model, but I’m going to learn with these new algorithms. These new algorithms, in addition to learning the knowledge, they’re also going to learn how to learn future things. So they’re learning two things at the same time. And then I think you would be able to learn new things without catastrophic forgetting.

    Alfred Lin: Is that the most radical thing that you’re trying to do in your company in terms of—from the current state of affairs to try to do these two things at the same time?

    Rich Sutton: Most radical thing.

    Khurram Javed: I think that’s ...

    Sonya Huang: This goes back to “I’m not crazy, everyone else is crazy.”

    Khurram Javed: Yeah, that’s perhaps not totally radical. There was a point in 2016 to 2018 where a lot of people were exploring these ideas quite a bit. They were doing it in a much more limited setting. So they would say we have a distribution of problems, and then in this specific case we’ll do it, whereas we want to do it from a single stream of experience. So our method should be more generally applicable. So I think many people have explored this, but no one has explored this in the general setting where the resulting algorithm would be applicable everywhere.

    Alfred Lin: So what would be the most radical thing that your company, your new company, is trying to do that other people are not doing?

    Rich Sutton: What’s the most ambitious thing? Remember, I don’t think I’m weird, so I don’t want to say it’s radical.

    Alfred Lin: All right, what’s the most ambitious thing?

    Rich Sutton: The most ambitious thing, I think, is to try to have the full spectrum of knowledge, both about the tiny things and about the big things. Like thinking about how you take an airplane from one city to another. That’s a very big—it’s like your space flight example, but it’s just kind of more commonsensical to think about, because many of us take airplanes, and all of us use abstractions in all kinds of our life. And even the squirrels use abstractions. So to have that spectrum of knowledge from the small to the big, and to treat it in a uniform way and to be able to have it self-maintaining. The big question is always, you have your knowledge-based system, and what keeps the knowledge in it correct? Well, what keeps the knowledge correct in a large language model is well, people did a lot of post-training, and they made sure it was correct and then they freeze it after that. So that’s what keeps it correct. But really, in our minds, we are always changing things, and yet something keeps it organized and coherent and settling back into a good place rather than drifting off into crazy land. That is, I think, our biggest ambition, to have a mind that is self-consistent and can keep training itself and making it coherent.

    Sonya Huang: I love that. Can I ask, it almost seems that it’s such an ambitious vision, and the idea that all these things can be unified into a single mind is so ambitious.

    Rich Sutton: It’s within reach. I think it’s within reach. Here it’s 2026, and our computers are so fast. Is it so ambitious that it’s out of reach? Or do we have already inklings of how all the steps can be done? And I think we have a vision and inklings. I don’t think it’s inappropriate.

    Alfred Lin: Your vision involves a trillion-parameter model with 20 watts. That seems pretty ambitious.

    Khurram Javed: That is ambitious. In some sense, with current technology, I would say it’s also impossible. Like, just storing a trillion parameters in memory would probably use more than 20 watts of energy with current memory technologies, but we are really thinking of okay, things are getting better, computation is getting cheaper, it is getting more energy efficient. So where would it be in five to ten years? And I think in five to ten years with the right algorithms, we can totally be in a world where this would be possible.

    Rich Sutton: So five to ten years is two orders of magnitude of Moore’s Law. It’s a standard improvement. If we double every 18 months, 10 years would give you two orders of magnitude. And so for Khurram’s statement to be plausible, then today you should be able to do it for—for what, 20 watts, two orders of magnitude, 2,000 watts. If you can do it with 2,000 watts today, then in 10 years you’ll be able to do it for 20 watts.

    Alfred Lin: You think you can do it for 2,000 watts? You have lots of people at research labs that have access to way more than that.

    Khurram Javed: Yeah, I think we can be more efficient than that even now with the right algorithm.

    Sonya Huang: If we can be more efficient than that, then why aren’t we? It’s not like people just want to spend all their money and spend all our money. [laughs]

    Rich Sutton: Sometimes it seems like they want to, doesn’t it? I think that’s how they show they’re real men, by using lots of energy.

    Khurram Javed: At least when I look at different research groups, I don’t even see anyone believing that it’s possible. And I think if you don’t believe in it, you’re just not going to work on the technical problems and work through them.

    Alfred Lin: Is it that it’s not possible or it’s that there’s so much waste in the system? Which one is it? Is there someone who knows how to do it efficiently?

    Khurram Javed: Yeah.

    Alfred Lin: And then there’s 10 times the number of people in the same lab doing all these other things. And so nine out of ten people are wasting ...

    Khurram Javed: In some sense, the way I think about it is that we are stuck in a local minima. So if we want to move towards these new kind of algorithms, it is almost impossible that things will not get worse before they get better. So when we start exploring these new directions, you’re not going to get state-of-the-art performance from day one, because it is a different paradigm. But that path leads to similar performance at a higher energy scale. And these big labs, they are so locked into a product that it is not possible for them to pursue a path where things get worse first.

    Alfred Lin: Because their current paradigm allows them to keep scaling, and this new paradigm, they have to take a bet, and they ...

    Khurram Javed: And they have to figure out some technical things that are difficult that we have thought about for many years. We know people who have thought about these things for many years. And when I talk to them, it makes sense that it’s doable, but you need to think about those challenges for a long period of time.

    Alfred Lin: So if everything goes right with Oak, what happens with the company? What kind of company are you building?

    Rich Sutton: If everything goes right, we implement the architecture, we can have genuine continual learning, and we can form abstractions so that we can do planning and reasoning. And we have sort of true intelligence. And then it’s hard to imagine just exactly what will happen by then. But I think ...

    Alfred Lin: Humans will become irrelevant?

    Rich Sutton: I don’t think that’s true at all.

    Alfred Lin: We don’t either.

    Rich Sutton: I think the world becomes exciting and even more exciting and interesting for humans. But in particular, I think there are the—you have to wonder about the large language models. They might be at risk when this eventually happens. I’m sure they’ll get a good run. They’ve already had a good run. They’ve been very successful. And let me say just for clarity that large language models are an amazing scientific breakthrough, a breakthrough in the skillful use of language by neural networks. It’s totally unanticipated. It was always a holdout for symbolic methods in language, and they have totally changed how that’s thought about now. So it’s a big breakthrough. So it’s frustrating to me that we have to—you don’t just celebrate that we’ve made this great progress in the subset of the problem of AI and enjoy that. Instead, it has to pretend to be all of AI. All of intelligence is not fluid, capable use of language. There’s so much more. It’s an important part. It’s like 20 percent or a quarter of intelligence. There’s more. We’re not done.

    Sonya Huang: Yeah. If everything goes right, are you imagining that there’s a single mind that can do everything from learn how to swing from tree branches to make a spaceship, to all these various things we’ve talked about today? Is it a single mind, and is it a single set of weights that can do all these things? Or is it ...

    Rich Sutton: It’s a single design.

    Sonya Huang: Okay.

    Rich Sutton: There’ll be many different—there’ll be many minds.

    Sonya Huang: Okay, so it’s a single design that reacts to different environments.

    Khurram Javed: And different versions of that mind would learn different things because they have different experiences. This sort of goes back to the big world hypothesis, that there are infinitely many things to learn, so one system cannot learn infinitely many things. And I think Rich already mentioned this, but if you have two of these systems, if you have two of the largest systems in the world, then it is trivial that they cannot model each other because they’re equally complex. So a single system would never be able to get to a point where it can learn everything. It would always be multiple systems that are learning from their own experience.

    Alfred Lin: You guys are hiring? What kind of people are you looking for?

    Khurram Javed: We are hiring. The initial team, most of it we already have in our mind. So these are people who have thought about these ideas in the past for a long time. And we are going to take a slightly different approach, because this is a different paradigm. It doesn’t make sense to become large very quickly, because in some sense everyone we hire has to come to see what we see. And not everyone sees that. So we’re going to start small, slowly grow to maybe a handful or two or three, and then go from there.

    Rich Sutton: We want to be super aligned.

    Khurram Javed: We want to be super aligned.

    Rich Sutton: So that we can be very productive working together and scaling the progress.

    Sonya Huang: Absolutely.

    Alfred Lin: Very, very cool.

    Sonya Huang: Wonderful. I love this conversation. Thank you for taking the time to share what you’re up to. You are an unusually deep thinker about where reinforcement learning and algorithmic design will go, and it was a true pleasure to get to explore it together with you today. So thank you.

    Rich Sutton: Thank you very much.

    Khurram Javed: Thank you for having us.

    Alfred Lin: It’s our pleasure.

    More Episodes

    Training Data

    /

    Joshua Meier & Matthew McPartlon, Chai Discovery

    Chai Discovery’s Bitter Lesson: Drug Design Is Another Scaling Problem

    Training Data

    /

    Matan Grinberg, Factory

    Factory’s Matan Grinberg: The Coming ‘Dark Factory’ Where Software Builds Itself