Everybody who has ever been interested in computer science knows Peter Norvig and Stuart Russell, the authors of Artificial Intelligence: A Modern Approach. Their book has introduced thousands of students to AI.

That’s my copy
I had the pleasure of sitting down with Peter Norvig at AI Conference for a conversation about what forty years in AI teaches you, and what can still catch you by surprise.
He is wonderfully comfortable sharing what he has had to reconsider over the years, and for me, there was an important lesson in that. He once doubted that millions of businesses could get online because so few people knew how to set up a web server. He understood the difficulty. What he hadn’t anticipated was that people would build tools to make it easy.
It made me wonder which of our assumptions about AI will look just as strange in a few years. Even his textbook has had to reverse its advice: a million rules once seemed impractical. A few editions later, it was “only a million.”
That willingness to reconsider runs through our conversation. We talked about why he thinks language models have further to go, what students should learn when AI can do their assignments, and why he would like to turn his textbook into something that changes every month.
There’s also a lovely moment when he suggests that students might not need to read the whole book anymore, but the LLMs definitely should. And when I asked whether he trusts the people leading AI companies, he was candid about that too. Read it!
Or – watch it.
We thank the AI Conference team and Jenna Dobkin for helping organize this interview
Subscribe to our YouTube channel, or listen to the interview on Spotify / Apple
This interview has been edited for clarity and length, and exchanges have been reordered. Separately labeled excerpts preserve their original context. Section headings and the introduction are editorial.
When the tools change
Ksenia: Do you feel like you can make any predictions?
Peter Norvig: Things will continue to change quickly.
I don’t have a very good track record for prediction. There was an ARPANET, and I was on that as a student at Berkeley. Then Gore and others had this idea that there was going to be an internet that was open to everybody. Anybody could sign up for it. Businesses would sign up for it. And I said, “That’ll never work, because there are only a couple hundred people in the world who really know how to set up a web server and do it right. How are ten million companies going to do that?”
What I didn’t realize is, when there’s a demand, people will make the software tools that make that process easy. You don’t have to be this Unix wizard to set up a network. You just buy this package from Netscape or Microsoft, answer the questions, and it works.
Maybe I’m the worst person to be asking, or us insiders who are saying, “I know all the details of how you build these language models.” You should be asking an end user, “What tools do you need to make your life work?” I think we’ll see these more powerful tools to do things that we used to think were hard, and then they’ll become easy.
From our discussion of what the first edition of the textbook failed to foresee:
I guess we didn’t foresee the scale. In the first edition, we said, here’s a description of the rules of chess. There are like ten lines, one statement for each piece. “For all pieces B, if B is a bishop, then B can move in this diagonal kind of way.”
Then we said, if you wanted to do it in propositional logic, you couldn’t have a rule for all bishops in all places. You’d have to have one rule that says, if a bishop is on this square and it wants to move to this square, then you’d have to have this and this. We said that would be totally impractical, because there would have to be like a million rules.
By the second or third edition, we’d say, “This is much more efficient, because there are only like a million rules.” In ’95, that didn’t fit into memory, but a couple of years later, it fit into memory. We have really good problem solvers, SAT solvers and other things.
The intelligent agent book
Ksenia: I think you define artificial intelligence as the way to understand rational agents. What are rational agents now?
Norvig: Over the last year, we’ve seen this re-emphasis on the idea of agents. We started it back in 1995. One of the things our publisher did for us was, we had the book and they were trying to say, “What’s unique about this?” We had agreed on the title, but then they also put a little banner that said “the intelligent agent book.”
At the time I thought, “This is kind of cheap marketing. I expect a sticker to be on my new and improved soap, but not on a serious textbook.” In retrospect, that was probably a good move, to say clearly that that’s the way we were going to describe what AI is.

Ksenia: How do you see agents now? I used your book a lot when I was writing the series about agentic workflows.
Norvig: I guess one of the things we really missed is, how do you build an agent? We did have a chapter on game theory and decision theory that talked about communities of agents and how they interact, but it was less of a focus than we’re seeing today. Right now, we’re seeing the main way you want to interact is to have a swarm of agents and figure out how they work with each other. Just having one is not enough.
Neural networks and what we build on top
From Norvig’s opening discussion of the textbook and the continuing role of older algorithms:
Stuart Russell, my co-author, and I had been teaching out of the other books. We thought the other books available were good, but it felt like the field had changed. Around ’85, there was a big push away from logic and towards probability. Judea Pearl was probably the leader of that. We were too constrained by Boolean logic, and AI should really be the study of uncertainty. Probability was the right way to do that. Machine learning rather than hand-coding was another change. The move towards big data was another change. We felt like the older books were missing those changes to the field.
Sometimes people ask me, “Do students still need to read the first half of your book?” The second half is on machine learning, and we covered, to some extent, transformers. But the first half is knowledge representation and logic. I guess my answer is, the students probably don’t have to read all of that, but the LLM should definitely read all of it.
If you look at AlphaGo, the reason it was able to beat the world champion at Go, decades before people thought that was going to be possible, is because it used neural nets in a clever way to encode what people knew about a board position. That would have been enough to demolish amateurs like me. But in order to beat the world champion, you have to take that and also put in Monte Carlo Tree Search, which is this traditional, old-school algorithm.
I think that’s still true. Going forward, we’ll have this mix of saying you need the neural nets to understand the complexity and uncertainty of the world. But you also need clever algorithms.
From his answer about the return of symbolic approaches:
There are some advocates saying we should go back to symbolic AI and maybe combine them in some way, having neurosymbolic AI. I agree with that, but I guess I think the current system of language models is probably approximately on the right path. I don’t think we need to tear it all down and start over again.
I think the core should be something like what we have now, a neural net with attention. On top of that, we should learn how to do logic. I think that’s the way humans evolved. We aren’t born knowing how to do modus ponens. We’re born knowing how to interact and see the world and move. Then we learn how to do logical reasoning on top of that.
From our later discussion of world models:
I guess I feel like language models are world models, but they could be better world models. Why is a language model a world model? Because it knows that “the cat chased the rat” is more likely than “the rat chased the cat.” It knows a lot about the world.
A language model obviously can’t do real-time 3D scene generation. That’s something useful.
But I guess I see a smaller distinction between a language model and a world model.
I look at Yann LeCun. He says models should be able to do A, B, C, and D. And I agree with that. I say, “Yeah, all those things are important. You’ve zeroed in on things that we don’t do so well today that we should do well in the future.” Then his conclusion is, we have to tear it all down and start over. Mine is, no, I think if we do small tweaks to what we have now, we’ll end up in the same place.
Three levels of feedback
Continuing that discussion, Norvig described how Google combined different kinds of feedback:
At Google, in the pre-AI days, we said, “We want lots of kinds of feedback.” We want to measure the clicks, and that gives us feedback that’s not very specific, but we have billions of data points. We can do experiments and change things and see if the clicks change. But that only tells you so much.
We didn’t call it reinforcement learning with human feedback, but we had human raters. We said, “Which is better, this set of search results or this set of search results?” We felt like we couldn’t have done it without that. It would have been incomplete to just have the click data.
Then we had a third level where we would take an individual person, put them in a room with a one-way mirror, record their every interaction and ask, “What are you thinking about now? Try to search for such and such. What search terms would you come up with? What do you think about this result page? What is this telling you?” And do this very in depth.
We had:
one level where there are billions of pieces of not very useful data,
an intermediate level where there are tens of thousands of more detailed data,
and another level where there are dozens of pieces of very detailed data.
It felt like it would have been incomplete without any one of those levels.
Power and the people building AI
Asked about AI’s positive potential and its risks, Norvig described a broader change in who can wield powerful technology:
Norvig: I’m concerned about that. My name’s on the cover of the book, so I don’t want, after the apocalypse, somebody to look and see, “This guy was responsible.”
I guess that’s true of every technology. I was watching the opening scene from the movie 2001. The bone comes down and the ape picks it up. We’ve now become a toolmaker, but one of the first things he does with the tool is bash his opponent on the head. I don’t know if that scene actually happened anywhere in real life, but we did decide as a species we’re going to be toolmakers. We did see that most tools have dual use and can be put to bad as well as good. I think that’s true of AI.
Even without AI, we’re seeing this push of enabling stronger technologies to be used by smaller groups of people. It used to be we felt like we were in this stable standoff: we had a couple of superpowers, and they had nuclear missiles and aircraft carriers, but they knew they couldn’t really go to war against each other because the outcome would be bad, and so it was stable.
Now you don’t need an aircraft carrier. You invest a couple tens of thousands of dollars in some drones, and you can exert your power in a way that wasn’t possible before. We’re seeing more of that, where smaller groups can be more powerful through technology. AI is adding to that. I think that does destabilize the world in some way, in that there are more threats to worry about.
Ksenia: Do you have conversations with the leaders of these companies? How do you see them building this? I think my question is, do you trust them leading this?
Norvig: I guess I’m worried. I have pretty good trust in Dario and Anthropic and the way they were incorporated. On the other hand, I worked with Sergey Brin daily, and I saw him change. I’m not as familiar with other people like Bezos, but from the outside it looks like they’ve changed as well.
I think trust can only last so long. Maybe there’s this corrupting power that, even if people start out with the best of intentions, sometimes they can be led down a bad path. I do worry about this concentration of wealth and power.
I’m somewhat optimistic that, if you had looked four years ago, at least I would have predicted all the attention was going to be on the biggest frontier labs. It was going to be this oligarchy, and just a small number of companies were going to be in charge of everything.
Now you’re seeing DeepSeek and Kimi and all these other [open-source models] saying, “We can have much smaller models, and they’re pretty good.” I’m optimistic that that means it won’t be a winner-take-most, that the power will be much more widely distributed. I think that’s a safer place to be in.
As we discussed whether open source could help, Norvig also raised its risks:
I guess there are also dangers of open source. I remember four or five years ago, Eric Schmidt was saying, “We can’t have open source because then some bad guys can download a model and ask it, ‘How do I build a pathogen that’ll kill everybody?’ And we won’t know.”
I was sympathetic to that point of view, but now everybody’s given up on that. Even Eric says, “It’s too late. The cat’s out of the bag. These open-source models are out there.” If we’re going to protect ourselves, we have to look at other ways of protecting ourselves.
Recursive self improvement and software engineering
On recursive self-improvement:
Norvig: To me as an engineer, that just means building more efficient systems, and that’s what we’ve been doing all along. I think that’s good. When we built optimizing compilers, that was recursive self-improvement, because that meant you could compile the compiler and it was better. I feel like that’s the same thing.
I guess what you want to worry about is, what are your software engineering practices? I think that’s true whether you have recursive self-improvement or not. We’re already starting to see these arguments: “Do we still need to do code reviews? If AI is writing the code, do the humans have to check it, or should we just let it go?”
Figuring out what the right level of control is, and how much these systems can do on their own, is the key to having safety. That’s true whether you have recursive self-improvement or not.
What students should learn
Ksenia: At the keynote, you listed a few things that need to change because of AI, and education is one of them. How do you see it changing?
Norvig: That was in the context of what software engineering is. We have to think, what are the skills that you want now?
What I’ve told my colleagues at Stanford and Berkeley is, “Don’t concentrate on the number of majors you serve. Instead, why don’t you try to optimize for the number of minors?” I think what’s more powerful is for somebody to say, “I want to be a physics major or a biology major or a sociology major. I want to know the AI and computing tools that I can use in my profession.”
That’s basically what mathematics has done. Very few people become professional mathematicians, but lots of people use calculus or statistics in their work. I think CS education should be more like that.
In many fields, we relied on essay writing to prove that you know the material. To some extent, you could always hire somebody to write your essay for you, but now the barriers are much lower.
One of the answers people have come up with is, rather than doing answers, maybe you should assign the students to come up with problems. If they can come up with a good problem, whether they solve it using AI or on their own, that exercise of coming up with a good problem was really the creative part, and that’s what they should be judged on.
Other people are saying, “I’m going to have oral exams or discussion rather than writing.” I think writing is a really useful tool. I spend a lot of time doing it. One of the weaknesses of writing is you don’t have to answer challenges. You say, “Here are my points, A, B, C, and D.” You can judge whether those points are all coherent. But in an essay, nobody can say, “What about X, Y, and Z?” In oral discussion, you can.
That works if you’re in a small liberal arts school and you have a class of a dozen students. It doesn’t work if you’re in a state school with a class of three hundred students. Are we going to break them up? Are we going to have AI tutors do some of the job? Are we going to do more with peer-to-peer learning, which I think is a very valuable thing?
Ksenia: I’m wondering how you approach the classes you teach.
The class is the human-centered AI class Norvig discussed teaching at Stanford.
Norvig: We try to make the class very interactive. It shouldn’t be just lectures. It should be students involved. We want them to ask questions and be involved in discussions, and we also break up the class and have them doing stuff.
Last time we taught it, for the first time, we said, “Starting now, everyone’s going to make an app.” About half were CS majors or other engineering majors who had lots of programming experience, but half weren’t. They were lawyers or sociologists or in other fields that didn’t have programming experience.
In that one class, everybody successfully wrote an app. Some said, “Wow, I had no idea you could do that.” I think this next time we teach it, we’ll have more of that. I like the idea of people doing it in teams and in an environment where others are commenting on what they’re doing, rather than having them go home and do it on their own.
A textbook that can keep changing
Ksenia: Have you thought about revising your book?
Norvig: Yeah, I’ve thought about it. Stuart’s take is he doesn’t want to do a new edition until he completely understands how deep learning works. I say, “Well, we’ll never have time to do that.” My take is, I would like to do it, saying what we know now, even though it’s incomplete. But I’m concerned about the format. I feel like our publisher and the whole publishing industry in general don’t offer the format I want.
They say, “You can write a new edition. You could do a new edition every year if you want, but we sell editions.” I say, “I want to sell a subscription that gets updated every month as new things come out.” They say, “We don’t know a way to do that.” I say, “I want there to be simulations and live code that people can run.” They say, “We can’t do that either. We could sell the PDF, but we can’t do that.”
I guess I’m waiting to say, is there a point where the publisher has those tools so I can offer that kind of thing, or where enough time has passed that I can write something new, none of the copyright is owned by the publisher, and I can do something on my own?
Ksenia: What is a forgotten idea that you think can contribute to current AI study and development?
Norvig: I hope the idea is that we’re building this technology to make people’s lives better. We should think about how we can help humanity, and not just how your company can get rich.
In a final follow-up about ideas worth returning to, Norvig pointed to Rich Sutton’s “The Bitter Lesson”:
Norvig: I like Rich Sutton’s take of the bitter lesson: the more you try to be clever and do something specific, it doesn’t work. But if you add more data, data gives you breadth, and then search gives you depth. Maybe it’s building systems that can make sense of data and can be searchable to a deep depth.
← Previous Interview: How Responsible AI Changes In The Agent Era / Related interview: When Will We Fully Trust AI to Lead? →





