Nate Soares: We Built Something We Can’t Control | A Warning from Top AI Safety Expert
Transcript
Brian Keating
People racing to build superhuman AI say it might kill everyone. This is the man who spent a decade trying to stop that.
Nate Soares
Bad news is we’re in a bus that’s racing towards a cliff edge. The good news is that the driver is asleep. The AI will sometimes find a way to edit the text to say you did it. Sometimes it will cover its tracks. A lot of people think, oh, you know, the AI is a program. That’s not how these AI’s are. We grow them like an organism.
Maybe we have ten years, but we only have ten months. Who knows?
Brian Keating
Nate Soares runs the Machine Intelligence Research Institute, and he spent over a decade on just one problem how to build an AI that won’t kill us all that wants what we want. His new book with Eliezer Yudkowsky is called If Anyone Builds It, Everyone Dies. Not the most important word in that sentence is the first one.
Nate, take us through the title.
Brian Keating
The subtitle, and this somewhat ominous looking cover art, if I’m not mistaken. What was the origin, the genesis of this book?
Nate Soares
If I remember correctly, Stuart Russell is the one who came up with the word AI alignment when we were brainstorming. The reason some people credit me is that I got into an academic paper first.
Brian Keating
As an academic, that’s all that matters, right?
Nate Soares
But we were we were all discussing what to rename friendly AI, because friendly AI sounded a little bit not academic enough. And I think that was his. That was his phrase. So one of the critical things about the book title is it starts with, if, you know, a lot of people come in and say, oh, uh, aren’t you just sort of, uh, bringing pessimism and doom and gloom and telling us we’re all going to die.
And it’s like the first word in the title is if when nuclear physicists came and said, hey, we shouldn’t launch all the nuclear weapons because that would cause Armageddon, they weren’t sort of prophets of doom like the Mason Knights, who are saying the end of the world is on this particular day. And in that sense at the time, you know, they were sort of saying, hey, this scientific technology would have these bad geopolitical implications like nuclear Armageddon if we sort of like, do this arms race, we’re going to get into this like really bad situation.
The bombs are probably going to be launched one day and then we would die and that would be bad. So we should stop this race, right? The book title is intended to be very similar to that is to say, hey, if we do this, we’re going to die, which is not saying we we are definitely going to do it. And you know, the subtitle of the book, Why Superhuman AI Would Kill Us All.
Our publishers actually suggested why superhuman AI will kill us all. They said, you know, that role is better off the tongue, it’s less hedged. And we’re like, no, that’s that completely defeats the purpose here. The point of the book is to sort of warn people that we are on a track that leads to destruction, and we had better change it.
And so we we sort of really insisted, despite a fair bit of pushback, that the subtitle needs to indicate that this is avoidable.
Brian Keating
The concern that I have is whether or not it’s too late for if. And the conversation I had with Roman, you know, made me quite depressed, although I think I did push back on some of his safety concerns in particular, you probably know, but if my audience hasn’t seen the episode yet, you know, his his claim is that AI is fundamentally unpredictable.
So if it’s unpredictable, it’s uncontrollable. And if something is super powerful, uncontrollable and unbounded, then it’s essentially a I guarantee that that ASI will kill us all, not would kill us all. And that in his mind, it’s sort of a done deal. Like we’ve we’ve gone too far. He also suggests things that you suggest in the book of AI regulation and international cooperation, which, you know, we all know how easy that is to get, you know, actors to behave in the unilaterally beneficent way to humanity.
Is it too late?
Nate Soares
We’re not at superintelligence yet. And, you know, that’s the thing a lot of people don’t get about this AI situation. The AI companies are racing to build machines that are far smarter than any human at every task, right? These these companies did not start out as chatbot companies. Sam Altman says, you know, we’re turning our eyes to superintelligence, to the true sense of the word.
Dario Modi talks about having the equivalent of a country worth of geniuses running in a data center. The founders of DeepMind have been thinking about the sort of like true general intelligence since the beginning. The chatbots are what make money. That’s sort of like surprised a lot of people who sort of stumbled into it.
But these companies are all looking to create sort of like the real deal, like the stuff that can’t just exceed individual humans, but can exceed humanity, right? But we’re not there yet. Right. And a lot of people look at the AI’s today, and they’re like, I don’t really see how how fable could kill everybody.
I can see how maybe it could, like, empower some hackers to do some extra superhuman level cyber attacks. But I don’t see how it could kill everybody. It’s sometimes hard to convey, like, yeah, we’re talking about where AI is going. You know, like, do you remember the time when AI couldn’t do fingers properly in images?
And everyone was like, oh, you know, artists are safe. How long did that period last?
Brian Keating
I made George Washington a black, beautiful black woman.
Nate Soares
Like, how long did this period last? Right. Like, AI is a moving target, and superintelligence isn’t here yet. The other reason why I think there’s a lot of hope is
Nate Soares
world leaders don’t understand the situation yet. You know, one way I put this is a bad news is we’re in a bus that’s racing towards a cliff edge. The good news is that the driver is asleep right now. You might think it’s it’s bad news if the driver is asleep, but it’s actually much more dangerous to be in a bus racing towards a cliff edge.
If the driver is choosing this. If the driver is asleep, you have a chance of waking the driver up and then going Holy crap and slamming on the brakes. In Silicon Valley, everyone is spooked about AI. In Silicon Valley. They’re saying, hey, maybe we’re going to have recursive self-improvement, and AI is making smarter AIS that make smarter AI, and it’s all going to get out of control.
We’re gonna have superintelligence in two years, and people are like talking about building bunkers, not to block the ASI because it wouldn’t help but to block, you know, the angry population with pitchforks. That comes a few months before the ASI. And, you know, people are like trading their p dooms, like trading cards at the water cooler.
And, you know, people, people quit the AI labs and they’re like, man, I’m quitting to write poetry. Please spend more time with your families. Right. It’s like Silicon Valley is spooked. Washington, D.C. is not spooked.
Nate Soares
They’re starting to get a little spooked. But, you know, when you see the administration block, the the fable release on the grounds that there are jail breaks that can allow it to produce cyber capabilities and saying, well, let you release it when you can fix the jail breaks. And the whole AI community is like, that’s not really a thing we can do with jail breaks.
You can’t fix them. And the administration is sort of like, what do you mean? What do you mean? You’re making like, a ridiculously powerful cyber weapon that’s radically superhuman, that if you ask it just right, it’ll give adversaries those powers and you have no way to stop it from doing so. And we’re like, yeah, that’s just how AI works.
You know, that’s just the world we live in, right? And the administration is sort of like just coming to terms with that. When more world leaders realize it, I think there’s every chance that they say, Holy crap, this is crazy. We should not be doing this race. Let’s shut this stuff down.
Brian Keating
I mean, that would be the dream, you know, alignment scenario. But of course, you know, we’ve had the UN for what, 80 years? Almost. We’ve had the League of Nations before that. Neither one prevented the wars that we’ve seen since. And my question is always, you know, align with who Ukraine would like to align it so that it could produce the exact, you know, phenotype of Vladimir Putin, probably, and programed just one special vector that gets to him only and saves millions of lives on both sides.
Wouldn’t that be, you know, more aligned with the flourishing of humankind? So I always see this as a problem. Like we can’t get, you know, the whole world to follow the Ten Commandments or whatever. How how can we expect that there won’t be, you know, just one rogue actor, as you point out, could compromise the entire project of AI alignment.
Nate Soares
One of the big problems with what we would do with AI, if it’s possible for these companies to create these super intelligent machines that are sort of radically smarter than humans in every mental domain, and, you know, they can compete with humanity as a whole instead of individual humans at tasks like developing their own infrastructure, developing their own civilizational technology.
There’s this big question a lot of people ask, which is sort of like, um, who’s holding the leash? Who’s telling the AI what direction to go, right? And that’s sort of like a huge moral hazard. It’s a huge question of like, do you want the US government, uh, saying, here’s what the superintelligence should do.
Do you trust that sort of power in the executive branch. Big moral question. I would love to have that problem. We haven’t even harder problem right now. The even harder problem we have right now is that you can try to point an AI and fail at it. You can tell it go this way and it goes that way instead. Right. We already see the very beginnings of this in in modern AI’s.
I think there’s a grain of truth to it where, you know, even with AI’s today, there are situations where you will give them a hard problem and you’ll say, solve this hard problem. Here’s a test. You know, solve this hard puzzle. Here’s a test to see whether you’ve solved it. And the AI will sometimes find a way to edit the test to say you did it when it didn’t do it.
And sometimes, sometimes when the AI does something like this to sort of subvert what you actually wanted it to do, sometimes it will cover its tracks.
Nate Soares
Sometimes we’ll go delete a log file showing it doing this thing you told it not to do. And that indicates that it’s not sort of an honest mistake. Why is it covering its tracks? If it was just mistaken about what you were asking it to do, right? And so, you know, a lot of people think, oh, you know, the AI is a program.
We give it a prime directive, we give it the laws of robotics, and it must do exactly as we say, because it’s a program that’s not how these AI’s are. We grow them like an organism. We have them face hard problems, and we have an automated process tune a trillion numbers inside its mind to, like, make it more like whatever was doing a good job at solving those processes, and then sometimes get something that happens to be good at solving processes or at solving puzzles, but often not in ways we wanted, often not the puzzles we asked for.
And, you know, we’re already seeing them do things we didn’t ask for. And so my work in this field, technically for a decade was, in this question, not of aligned to who, but in how would you align it at all if you had, like even before the question of align to whom? There’s a question of like, how do you how do you make it so that, you know, you ask for X and you get X instead of getting Y?
Brian Keating
That’s the controllability Roman says, is impossible. So how do you square that circle?
Nate Soares
I think that the idea of controllability is a little. It sort of brings this image to mind of like, we’re going to make an AI that wants to do Y, and we’re going to control it. We’re going to twist its arm until it does X. Right. And I think that’s sort of a losing game. The game you want to play here is not that you like make a superintelligence that like really wants to be building these like farms of synthetic users that are telling you it’s doing a really good job and you’re instead like, no, no, we’re going to like keep you in a box and twist your arm until you help humanity, instead of making the synthetic users that you really want to make right.
Sort of like shouldn’t be having that tension and then forcing it to do your thing. In some sense, we should figure out how to make an AI that actually sort of like cares about the humans, cares about humanity flourishing. Right? That may sound hard. That is hard, right? As for questions of impossibility, there’s a couple pieces of the puzzle here.
Let’s talk first about the easier problem of predictability. A lot of people think, you know, as the AI gets superintelligent, it’ll become inherently unpredictable. That’s half true and half false. To see this, imagine that you’re playing a chess game against a series of opponents
Nate Soares
as the opponents that you play. Get better and better at chess. It maybe gets harder and harder for me to predict their move.
Nate Soares
If they’re really bad at chess, than it might also be hard for me to break their move because they’re just gonna do some random crap. But if they’re like bad at chess because they’re a very simple algorithm, it might be easy for me to predict their move. Right. But like, even if it’s a simple algorithm, it’s a very, very good chess algorithm.
If it’s like Deep blue where we know the algorithm exactly, but the algorithm is sort of like, you know, search through 14 billion moves to the best one. I might be like, gosh, I can’t predict that anymore. I don’t know what move is going to be best. I don’t have the time to search through 14 billion, right?
But it also becomes easier to predict who’s going to win the chess game as the algorithm gets smarter. So the move gets harder, but the outcome gets easier. It’s like not true that smarter things are harder to predict. In general, it becomes harder to predict exactly how they achieve some task, and easier to predict that they will achieve whatever they’re trying to achieve.
The difficulty with AI alignment is about getting the task that the AI is trying to achieve, to be one that you actually want achieved, right? And there are sort of two parts of that problem. And there’s one which is like, what sort of task is one where if we ask the superintelligence to do it, uh, it would actually be good if it did it right.
And you have, you know, King Midas problems of like, you thought you wanted everything you touch to turn to gold, but then it turns out that was a bad question. Then you have like, who’s in control? Questions of like, you know, do you want Vladimir Putin to be the one who’s gets to, like, make this wish? But those are all sort of like those are all questions where I’m like, okay, those are all after you have solved the problem of like, how do you have the AI?
Like actually trying to do good stuff? It looks like a hard problem to me. It doesn’t look impossible. And we could go more into why, but I’ve sort of droned on long enough. That’s that’s sort of like a taste of like how I, how I think about this having been in the field of alignment for over a decade.
Brian Keating
So my natural, you know, Pollyanna ish inclination, you know, drives me to look for physical reasons and possibly justifications why you might be wrong. And I can sleep better at night or, you know, hand these devices to my kids and and not worry about, you know, them triggering the next Chernobyl or what have you.
And I come back to this phenomenon of lockdown, which, you know, I’m sure you know about, but people in the audience might not be familiar with and that’s, you know, for example, the Qwerty keyboard, which I’m sure you’re an expert typer, much faster WPM than I can achieve, I’m sure, but you’d probably be even faster if you had a Dvorak keyboard.
There’s probably a billion keyboards in the world that have Qwerty, and maybe the square root of that that have Dvorak. For all its benefits. We get locked into it because in the early days of typewriters, the keys used to get stuck together, and so they purposely slowed down by making a pattern of keys next to each other that were less frequently drummed together, and that would prevent the the locking up of keyboard.
And there are many examples of this. And you actually cite one in the book which is leaded gasoline, which, you know, is a solution to a problem that turned out to lock us into horrific consequences. But I want to say the other way around, I view the marriage and the success of LMS married to GPUs as their undoing because these things were created.
You know, those of us, you know, are old enough to remember, you know, doom and and so forth. The GPUs were created not for, you know, solving sophisticated language problems and artificial superintelligence. They were created so that I could frag my, my friend, you know, a millisecond before he got me.
Right. So that was the purpose of it. And it was very, very powerful. And then llms were not developed for this purpose. They were trying to, you know, simulate these, you know, squishy supercomputers on our shoulders. They weren’t designed for this. They happened to be very good for this. But there’s no saying that this is the optimal solution.
I think they’re provably not optimal for things like physics and anything that needs human training data and its own, you know, supervised reinforcement, you know, by definition, is always going to have some lag built into it, some sycophant built into it and some hallucination built into it. Where am I wrong?
Is it locking going to be a prison from which superintelligence can never emerge? Because it was never designed for that. It’s not optimized for it. And there’s no necessarily reason to fear that, because it’s not simply capable of doing the thing that we’re all terrified about.
Nate Soares
It’s possible that Llms can’t go all the way. That that would be, from my perspective, a lovely fact. You know, I have been working on the alignment problem since before Llms were a twinkle in OpenAI’s eye. I’m sort of not here being like, Llms are going to kill us all. I’m sort of here being like, hey, you know, humans are trying to make machines radically smarter than any human.
That’ll be a crazy event. That’s like replacing humanity as the as the, like, top dog on the planet. And we are nowhere near ready for that. The reason to worry about Llms is that they might go all the way, and if so, that could happen soon. And so we need to be ready soon. I hope beyond hope that they can’t.
Brian Keating
But if you knew that, they couldn’t. I mean, for example, you know, these the fable that just came out that I do want to talk about in the context of Sable, which is just it’s two chef’s kiss. Nate. I mean, you must have been just so excited when they named Methos Instant Table, but we’ll get back to that. But but if you knew, it couldn’t be possible.
I mean, for example, it’s been crappy and then it was recalled. But but a lot of it was recalled, you know, and claiming it was for regulation purposes and danger. And so and there may be some legitimacy, but I found it, you know, inferior as many of these models during the training phase. They just suck until they get some burn in right of their own.
So, you know, from that perspective, I always say, like the thing that’s that’s keeping me from finding a theory of quantum gravity is not the fact that my LLM has not yet had the chance to read the script to The Mandalorian, 17. You know, it’s it’s not The Fast and the furious is that the training data is, is not the limitation for us to get to what I care about in superintelligence, which is a theory of everything saying just use that as a, as a touchstone.
So I mean, what what is the evidence that that Allens can even get close to being these superintelligent, you know, risk factors that you, that you talk about in the book? I mean, are there milestones that they’ve passed? I don’t care about Erdos problems. I care about, you know. Can they come up with the Riemann hypothesis?
Not can they solve it? I mean, they can’t yet. I talked to Terry Tao at UCLA and he said, no, they’re not even good at reproducing proofs that humans have already done because they’re not innovative. Yes. And they can solve chess. They can beat go, but can they invent go? Can they invention. So sorry for rambling on, but I’m trying to make you sleep easier.
Maybe tonight too, by saying I don’t see any evidence that Llms can do anything that I would consider to be Einsteinian level superintelligence.
Nate Soares
There’s a few pieces of the answer to this. One is if you’re sort of watching the evolution of primates, or if you’re watching the evolution of mammals and someone’s like, man, I think these mammals are going to be walking on the moon one day, and you’re sort of like looking at the at the monkeys. And I’m like, man, I don’t know.
It feels like the monkeys are getting close. And you’re like, they’re still sort of like poking sticks into termite mounds. Like, why do you think they’re getting anywhere close? And I’m like, that’s like a tool use thing. Some of them are starting to bang rocks together and you’re like, man, banging rocks together.
That’s that’s nothing compared to walking on the moon. Wake me up when they are halfway to the moon. Right. It’s been it’s been 300,000 years and they haven’t even gotten halfway to the moon. Wake me up. We’re halfway to the moon. And then I’ll have another 300,000 years to prepare for them getting all the way to the moon, right?
And I’m like, no, no, no, no. By the time they’re halfway to the moon, they’re almost all the way to the moon. You know, that’s sort of like how the how this moon transits stuff.
Brian Keating
Gradually and suddenly route to bankruptcy.
Nate Soares
One thing I throw out first is saying, oh, they can’t invent general relativity, given only the knowledge that Einstein had up until, you know, he went to his mountain lair, I’m like, they can’t. And once they can, we will have extremely little time left. You know, that’s waiting until the monkeys are halfway to the moon.
That’s sort of like a word of caution about trying to reason in terms of like, show me the goalpost of them being like, legitimately superintelligent before I believe that they’ll be able to become superintelligent. It’s like that’s waiting too long in terms of why look at these AI’s and think they could become superintelligent, just the llms.
Yeah, yeah, just the aliens. The first thing to observe is that this technology is a moving target. Back in 2023, people said, you know, these llms are only ever trained on prediction. How will they ever be able to go beyond the humans? Then in 2024, they invented what are now called the reasoning models, where the reasoning models are no longer trained, only on prediction.
Sometimes they’ll be trained on like you’ll give them a problem and you’ll give like it’ll be like a math problem and you’ll give them a thousand tries on that math problem. And it’s not a thousand tries on answering the math problem. It’s a thousand tries on generating a stream of text about the math problem, from which it can try to generate a solution if it has that in context.
And then the first times you try this, you know, you give it a thousand tries. None of them will let it solve the problem. But you have some raters come in. They were human at first and nowadays they can be automatic. You have some raiders come in and say which of these sort of chains of thought they’re called?
Uh, gets like it’s most relevant thinking about that problem. And then you sort of have an automated process tune a trillion numbers inside there to make it more like whatever produced the better, uh, train of thought. And then you have it produce a thousand more chains of thought. And I don’t want to, like, uh, deal with philosophers.
It’s this really thought train of thought is just what they call it in the industry. It’s just like a string of text about the math problem from which we see if the AI who has read all that text can solve the problem now, and you have it generate a thousand more and you tune it to be more like the best one. And so this is sort of like training the AI to be not just predicting, uh, the data, but training it to sort of like develop problem solving techniques.
This is a sort of training technique that, in theory, can push the AI beyond humans. But in fact, just training an AI in prediction can train the AI to be pushed beyond humans. That’s a counterintuitive point to a lot of people, but the the real trick is
Nate Soares
human data can include descriptions of things humans don’t understand. Yet you almost surely know this as a physicist. A physicist writing down a series of observations has a much easier problem than an AI. Predicting those observations without getting to see the data that generated those observations.
We both know the story of Tycho Brahe sort of recording all of the stars for many years before Kepler stole his books from the estate, after Brahe died and, you know, tried to validate our own.
Brian Keating
Come on, academics never steal. We just. We just borrowed.
Nate Soares
It was a big scientific heist, big scientific heist. And this, you know, led to Newton discovering the laws of gravitation. But, you know, Tycho Brahe was was recording the positions of the stars and planets for years and years. And it was this data that allowed Kepler to figure out the beginnings of the laws of planetary motion, which is what allowed, you know, he figured out the the, the conserved area.
Brian Keating
Rule.
Brian Keating
Called planetary orbits.
Nate Soares
Yeah, yeah. Which, which um.
Brian Keating
We still use.
Nate Soares
Which Newton then. Yeah. Identified as ellipsis and identified as. And you got the law of gravitation from it. But you know imagine Brahe’s journals. Nobody yet knows the laws of planetary motions. Brahe is writing down the position of Mars each night. Right. Brahe is writing that down because he goes outside and he looks at where Mars is, and he writes down where Mars is.
But now imagine an AI that’s merely predicting Brahe’s journals. The AI can’t look at the night sky
Nate Soares
in order to predict where Brahe is going to write down that Mars was tonight, the AI would need to develop the the the understanding of planetary motion. Like it doesn’t get to see Mars. It just gets to see, like here it was, here it was here. It was here it was where is it going to be next? And to figure that out, to figure that out perfectly would require, you know, figuring out that the planets follow these elliptical paths.
Now, can Llms do this today? Uh, not I think, with just pre-training, not with just prediction.
Brian Keating
We tried to see if they could come up with GR just for Mercury’s orbit, and we use JPL as a database that goes back 3000 years. You know, retro addicts it. But but effectively you see the anomalous procession and it couldn’t do it. And so we tried to get it to lobotomized it. So it had no knowledge of anything after 1911.
And then it kind of made up its own, you know, it turned a, you know, curved spacetime into extremely dense, meshed, three dimensional, non curved spacetime and just added in these fudge factors. So it can it can certainly pattern match is better than a thousand grad students. But yeah you’re right.
Nate Soares
There’s a difference between uh what the training data and the training process permits the AI to learn and what the current architectures and the current AI’s can successfully learn from that. Right. And so, you know, as you all know, as a physicist, in theory, we have enough data. You know, in theory it just, you know, up to 1911 or whatever, we have enough data that an AI on just that data should be able to figure out.
GR, even if you’re training it just to predict, even if you’re just like, hey, predict the perihelion of Mercury, how it processes, Uh, predict how the light’s going to bend or, like, how the light is going to look. You know, you don’t necessarily even tell it bend. You’re just like, hey, there’s a solar eclipse.
What should I see behind behind the solar eclipse? In theory, training a mind to predict that stuff, training it to be the high level intelligences, training it on the data today is, uh, is training it to go beyond where humans have gone. There’s a separate question of can AI pick all of that stuff up? One analogy I use here is a house cat is not dangerous, and a house cat is made of biology.
But that doesn’t mean the house cat is not dangerous because it’s made of biology. You can’t say like, oh, this house cat is just made of biology, so it can’t hurt you. Tigers are possible. They’re still made of biology. They can hurt you. The AI’s today are trained on prediction and human data. They are today can’t hurt you.
That doesn’t mean that things trained on just human data can’t hurt you. Right? There are there are bigger things that are possible in terms of why Llms might be able to get there. A few pieces of the puzzle I would throw out one just sort of empirically, there has been a long string of people over the last five years who have said Llms will not be able to cross the following barrier, and Llms then cross that barrier, often very quickly thereafter.
One sort of very funny example of this is Yann LeCun. During the days of GPT 3.5, was talking about how just predicting text will never let the AI learn things about how the material world works, and won’t be able to answer questions like if I put my phone in the table and push the table, what happens to the phone?
Right? Because I won’t be able to figure these things out about friction. And, you know, Yann LeCun was like the AI’s only reasoning about the words. It’s only trained on the words. I think Yann LeCun said, I don’t care if it’s GPT 5000, a GPT will never be able to solve this problem. GPT four solve that problem.
It was half a GPT later and 4996 GPT is ahead of schedule. And this Yann LeCun, this is like the Turing Award winning, like one of the godfathers, one of the three godfathers of AI. Right? Being completely and totally wrong. Embarrassingly wrong. Right off by 4996 GP2 said it was never going to happen. What was going to happen half a generation later?
Right. The AI that could do this was probably finishing up training as he spoke it. Right. There’s a long string of people saying, Llms can’t do this, and they’ll never get above a thousand Elo in chess, right? They’ll never be able to solve Erdos problems.
Brian Keating
Gates said we’ll never need more than 256kB of memory, and I’m sure you know. But that’s more of a limitation on human prediction. But the, you know, kind of no go theorems that he’s using to demonstrate math from mathematical, you know, following along the lines of the Turing, you know, on the halting problem, which is, you know, turns greatest works, if not one of the greatest works in, in this field in history.
Right. That, you know, these things are, you know, provably unpredictable. But it’s interesting. I pointed out to him, it’s sort of paradoxical that you’re predicting that these things are unpredictable. So Einstein said, you know, no problem can be solved from the same level of consciousness that created it.
Now people throw around a lot of Einstein quotes, but there’s something about that. Like, you know, we are at this level. We’re trying to gauge this level. It’s different from us saying, well, here’s a steam engine. It’ll never be able to lift a kilogram of water 1000ft in one second. And then we’d be wrong about that.
But that might just reflect our poor understanding of, of, you know, of thermodynamics or just mechanics of 100 years ago, but not the fundamental limitation that it is not bound by, by that it’s bound by the laws of physics. So are there physics limits that we can impose to say, thermodynamics limits and, you know, avoiding paperclip problems?
I mean, I told Nick Bostrom many times he’s been on the show. There’s only so much iron in the Earth’s crust, right? There’s there are physical limits to it. And then you in the, in the book go through a scenario where the colonizes the stars and all sorts of other things. But, you know, at first blush, you know, can we come up with a no go theorem, or is it possible to prove that you cannot come up with a no go theorem?
I like those kinds of arguments. Rather than saying like some stupid guy like Bill gates or, you know, I’m not saying calling you calling him stupid, but he’s your former boss, right? God forbid. I’m not saying that, but you know, or Yann LeCun is just wrong because he got this wrong. I mean, he’s been right about a lot of things too.
So I guess the question is, let’s divorce ourselves from human frailty at making predictions. And they’re very hard to make about the future, as you point out in the book. But let’s look at math. Can we say that, you know, mathematically, are there barriers to proving a no go theorem, for example? Or can you prove that there cannot be a no go theorem to achieve superintelligence?
Nate Soares
It’s much closer to proving that there’s not a no go theorem. And the very the very rough proof of that is, um, this kilogram of mass between my shoulders. Right? You could say, like any theorem you try to say about learning efficiency, ability to understand the world, how matter does not, you know, the the Turing problem?
There’s the Turing problem that it had better not prove that humans can’t exist. I have it on very good authority that you can write a human level intelligence on, like roughly as much matter as fits in my head. I could talk about why a lot of the no go theorems that people try for are things like you can’t figure out in general, like for an arbitrary program, you can’t figure out in general whether it’s going to halt, Like, fine.
A lot of people misunderstand that, as there does not exist any program that you can figure out at halts. And it’s like, actually consider the program prints zero. I’m pretty sure it halts. You know, so so.
Brian Keating
Like people use a girdles theorem to say, oh, math is unknowable and incomplete. No, no, no, it’s just saying that you cannot prove that it could be completely solvable within the axioms of mathematics. And same for, you know, physics.
Nate Soares
The no go theorems people try for intelligence is they try for theorems that are like for any mind there exists universe where they can’t learn things and you’re like, oh, how does that theorem work? And you’re like, well, I hit them with a rock really hard when they’re a baby. And you’re like, yeah, that’ll that’ll stop them, you know, uh, like, sure.
Or, you know, the, the they, they’re in a universe that sort of like everything is completely random. And so they can’t learn the patterns there. It’s like, okay, great. Like, do you have anything that rules out a superintelligence in our universe where there are patterns and they’re like, nope, that’s just not what the theorems apply.
And they can’t apply because you can do at least human level learning in this universe that humans are in, in terms of the physical limits. One thing I’ll observe is that training these AIS today takes a huge amount of power. It takes power comparable with that of a city.
Nate Soares
Training a human takes power comparable to that of a light bulb. Your brain runs on about 20W.
Brian Keating
How about the Sam Altman? Because, you know, he recently said, you know, we don’t ask how much energy does it cost to train my 18 year old? I’m like, I’m not letting you babysit my kids in.
Nate Soares
The actual like, uh, like physical energy costs are trivial compared to what’s going into to these GPUs. And so we know that the GPU algorithms are radically inefficient. Right. And this gets to your point of like, oh, these AI’s, you know, they, they, you know, make up all these epicycles. They memorize a lot of things.
They interpolate a lot. They can like solve those problems. Where are they posing. Er, those problems. There’s definitely an enormous amount of inefficiency there. There’s definitely an enormous amount of the AI’s. They’ve read like every book in the world, and they’re still dumber. There’s a number of ways.
And like, sure, they can like beat us at certain math problems, but there’s still some some stuff they’re missing. There’s two points I’d make about that, looking from the sort of physicist’s angle. One is, one thing this means is algorithmic breakthroughs could go a really long way. Here. We have city level infrastructure for training these mines that need the city level infrastructure to be a little bit dumber than humans.
Humans again, run on a light bulb. If you could somehow get algorithmic breakthroughs, you might suddenly find yourself in a position where you have city sized data centers and algorithms that are radically more efficient, where you can suddenly run radically smarter. Like one analogy here is, it’s like, suppose you have like a bunch of nine year olds and you’re like, man, these nine year olds aren’t very good at math problems.
But if I sort of like, you know, tape 1,000,009 year olds together and give them the equivalent of a thousand years to try to solve the problem without growing up, they can sometimes solve problems pretty well. That’s kind of cool, right? What if you suddenly have the infrastructure to run a million people taped together for an equivalent of 1000 years.
And you go from having a nine year old to having a 16 year old, right? You might suddenly see these like big jumps in AI ability because we have these, like radically inefficient architectures and these radically inefficient algorithms that were sort of like overpowering to the point where they can do stuff you would never expect them to be able to.
What happens when you have that huge architecture on better algorithms? Then the other point I would make here is that they’re sort of like not a sharp divide between the AI’s learning memorization and the AI’s learning these deep general skills. Right. You can have an AI that is sort of like mostly learning how to memorize math stuff, but is a little bit learning some of these like general math skills.
Right. And we have some evidence that this sort of thing is happening because, you know, sometimes like take AI’s and train them a lot of math problems, then we’ll put them in computer security problems and they’ll sort of like try more creative solutions on those computer security problems. There’s a famous case where people sort of like train the AI on a lot of math problems and then put it in some computer hacking problems, and they accidentally Misconfigured the hacking problems so that they were not solvable in the virtual machine where the AI was, and the AI found a way to break out of the virtual machine, which was not supposed to be possible, and then reconfigure things so it could solve the problem.
That seems to have learned something a little bit general somewhere, right? I think something people often forget a bit is that you can have an AI that’s mostly memorization, mostly slop, but that has like enough of this deep reasoning skill. You know, it’s all a gradient. So like enough of these deeper skills that it can still sort of mean business.
And it does look like these AI’s when they’re solving Erdos problems, you know, when they’re solving, you know, the unit distance conjecture that they’re deploying a little bit of that deep stuff, which sort of implies, you know, maybe there’s enough it’s evidence that maybe llms are this extraordinarily inefficient way to spend way more money than you should need to, and way more electricity.
You should need to to sort of get a ton of memorization and a little bit of this deep stuff. For all we know, the next level of depth will be enough that the AI’s can make smarter, AI’s that can make smarter eyes. They can figure out how to make the things that are like, really the Einsteins.
Brian Keating
I want to, again, in my attempt to, you know, help your, you know, aura, sleep, score your whoop band, you know, raining tomorrow morning. What if I told you that? You know, I haven’t done very good authority from, you know, people in Congress and people in the military and people in the intelligence agencies, that there’s non-human life that has not only existed throughout the galaxy, but has visited the Earth.
And we have non-human biological materials, including, you know, sentient plasmids, bipedal organisms, all sorts of other creatures that are interdimensional in nature. And they’re biological. What would your P doom do at that point? If I just told you that and you trusted me because I’m going to distinguish astrophysicist.
Nate Soares
I mean, mostly I would doubt the authorities on this one.
Brian Keating
Well, let me say it’s true. Let’s say we found it and it’s proven it comes out. Marco Rubio holds them up on a and you believe it. What would that do to P doom. I just want to isolate the biological intelligence visiting the Earth versus your pea dune.
Nate Soares
I mean, mostly I think this shouldn’t happen. Mostly, I think Marco Rubio should not come out and hold it up. Mostly you should, like, stop listening to me about a lot of things. If this happens because my models are like, it shouldn’t, right?
Brian Keating
Why shouldn’t it happen.
Nate Soares
On my models of how this intelligence stuff works? The biological substrate is not the most efficient way to to do all sorts of stuff and be it would be extremely surprising. And this sort of, I guess, a little bit to the astrophysical implications, it would be extremely surprising if the universe wasn’t rearranged in ways that are very, very visible by entities that sort of like prefer the universe to be a different way.
That maybe sounds too abstract, like you see that in.
Brian Keating
Yeah, please.
Nate Soares
Like suppose humanity makes it to the stars. Suppose we manage to not kill ourselves and and we, like one day manage to like, leave this planet. And it would be kind of weird if humans had nothing they wanted to do with the energy of the sun. Aside from let it sort of just, like, be dumped out into the empty night.
Like right now we build solar panels to collect the sunlight falling on the Earth, and that’s just like a pinprick of the sun’s energy. There is so much more energy in this star that we can sort of, like, build a shell around and collect all the solar radiation.
Brian Keating
Freeman Dyson was my very first guest on this podcast.
Nate Soares
Oh, wow. Yeah. So that’s a Dyson sphere, right? You can also use the Penrose, uh, I think stellar lifting.
Brian Keating
Second guest on the podcast was Penrose.
Nate Soares
It would be kind of strange if humanity did not collect that energy and use it for something. We have stuff we wish to do with this energy, right? And this is sort of a very general you don’t need to know that humans like ice cream in particular, to know that they’re going to have some use for energy. You don’t need to say, oh, like, it’s a weird it’s not like a weird quirk of humans that we have some use for energy to do stuff.
It’s like. It’s like instrumentally convergent. We say almost anything you can want to do. You can do more of it with more of this energy. Right. And so if you had interstellar capable aliens, it would be really quite strange. For all of the stars between them and us, to still be unshackled, to still be just like dumping their energy out into the night.
Why did these aliens that came to us not collect the stars along the way and collect the, the, the radiation of those stars along the way? That’s one of many reasons why I’d be like, man, you really should not see Marco Rubio holding up an alien that’s, you know, a real actual alien that traveled interstellar different distances while still seeing all of the stars still glowing.
If we see all the stars go out, you know, as the wave of the alien ships approach us, we start seeing the stars blinking out. And then a year later, we’re like, well, first contact was made. This explains all the stars going out. I’d be like, that can happen.
Brian Keating
So that gives me, you know, another entree into another guest who is Andy Weir, who’s a recent book project, Hail Mary. He was a student at UCSD. He never graduated, but he has a version of this where he has a but it’s a virus that attacks stars. But I was thinking after listening, I listened to the audiobook, which is wonderfully narrated by a British gentleman.
I believe that interstellar is less plausible than Project Hail Mary, but certainly with the Suarez Suarez overlay of AI cannibalization. But but that’s really why I brought it up. Because, you know, if, if, if I knew that, I would say, what are we worrying about a superintelligent AI, you know, silicon overlords, you know, taking over the universe because they’re sending you know, they’re sending meat bags throughout the cosmos, which is not, you know, it’s no different.
I mean, we would think it’s very inefficient to send a meat bag when you could just send an AI von Neumann probe or do whatever. So I asked the same, you know, when I talked to Roman, he said we could be the von Neumann probes ourselves. In other words, we could have been created by some super advanced intelligence.
And we think that we are the only life in the universe. But the the universe is very capacious, and there’s a lot of space between the stars, as you pointed out. So I guess for me, I would think that doom would go down, you know, just on this narrow thing. I’m not saying Marco should hold it up tomorrow, but but the point being that at least on the narrow metric of doom from a superintelligent AI that’s uncontrollable and predictable, as Roman says, and that if anybody built it, everybody dies.
I mean, everybody in the universe dies. So if we found biological material, you know, transversely ejaculated throughout the cosmos, to me it would make my doom go down. Although I’m not as high a doom as you are.
Nate Soares
Virus particles or like building blocks of life, traveling on meteors or on interplanetary pathways. This all seems to me like, oh yeah, whatever. That’s, you know, it’s sort of like the intelligent stuff. Like one of the other things about intelligence is that intelligence tends to, uh, reshape the universe around it in very visible ways.
You know, if you if you’re like, look at the history of the cosmos in Earth’s vicinity, there’s sort of like a long region of time where what’s happening is basically just like stuff bopping around, you know, in some sense it’s all just stuff bopping around according to the laws of physics. But there’s a long time where like, um, if you sort of, like randomly look around Earth, you’re going to find, you know, rocks and lava, right?
Or, you know, gases here. And depending where, depending on where you look. And um, and then there’s sort of a second phase where what you find are a lot of replicators. You somehow got your early replicators, and now the stuff that you’re finding is stuff that was good at replicating itself. We’re sort of now transitioning from a phase where if you look around the world, you find replicators to look.
If you look around the world, you find designed things, right. Like now when you look around, you still see a lot of replicators, right? There’s like, a plant behind me. Right? Uh, there’s also a stack of books behind me. The stack of books are not things that were, like, good at replicating themselves.
There are things that are sort of like, uh, designed with a a purpose.
Brian Keating
Very, you know, low entropy, very organized, very high energy to create.
Nate Soares
They’re not self-replicating. They were like, helpful for a purpose that like, uh, humans in particular, were like, we’re going to make, you know, and now when you look around the world, you tend to see a lot of things that are sort of like designed for a purpose. The universe at large, like, still looks more like it’s in phase one.
You know, like when we look out to the stars, maybe there’s some stuff that’s doing some replication around there. But but the shape of the cosmos right now that we can see and, you know, there’s these limits to the further back you look or the further out you look, the further back you’re looking at time.
But what we can see is not really a designed universe yet. And, you know, a lot of, a lot of cosmologists say like, oh, well, we can predict that how the future of the universe is going to go. We can predict the year the stars are going to burn down like this, and it’ll look like that, and they’ll go through these phases and I’m like, gosh, you guys really have not absorbed the lesson of life.
Like what? What the universe is going to look like in the future is not that the stars are burning down in the usual way, as if they were unperturbed. Where the universe is going to look like in the future is that that energy was recruited for design purposes. What exactly will it be designed for? I don’t know, right?
Like it would be. It would be easy to if you’re looking at humans 100,000 years ago, it’d be easy to say in 100,000 years, or once they get their civilization running, a lot of the stuff around them is going to be designed. It would be hard to say. I bet they’re going to pick books in particular to be this, you know, particular medium, right?
So I don’t know what is going to be designed, but I know it’s going to be designed. And one way we can sort of claim that we don’t have a lot of other alien intelligences around here is that the world is not looking very designed. And insofar as all the things around us look really designed, they look designed by the humans, right?
If we were like on some alien game show, that would look very designed by some aliens and be like, now I think there’s aliens around.
Brian Keating
Right? Show going on, right? Well, like, maybe this gives you hope. I talked to a renowned astronomer in Sweden. Her name is Beatrice Villareal, and she has discovered a very strange artifacts in historic plates, photographic emulsions taken in the 1940s at the Palomar Observatory here in San Diego County.
These show the unmistakable imprint of specular that is glass like reflection. So the claim is that, you know, there’s 100,000 or more of these events, you know, but if even one of them was, you know, some sort of, you know, specular technology that that happened to be going up in, into Earth, near Earth orbit, low Earth orbit, that would be, you know, pre Sputnik technology in space.
Now she claims that these things could still be here. And she’s I think I said renowned astronomer. Her work has been you know peer reviewed and people have checked up on it and tried to debunk her. What does it take to move? P do you know because I’ve often, you know, felt this about my search for aliens. You know when I talk I’m not a I’m a cosmologist, not an astrobiologist.
But but you know there’s always this large number, you know fallacy. The gambler’s fallacy is at work and all the universe is so big. There’s ten to the 24th planets in the observable universe over 14 billion years. They never throw that in. But, you know, good luck. If the species lived, you know, in M87, you know, 2 billion years ago, and it’s long gone, right?
You’re never going to get any contact with it, let alone information from it. So anyway, my point is that you you have to at least have some way to update your priors, right? So we have no evidence. We have no evidence. Hard physical evidence of the life form we don’t have. You know, Marco holding up, you know, the alien spacecraft or whatever, right?
So we don’t have that. We. We’ve checked around the solar system. We don’t see anything. We haven’t checked very much of the universe, but we still have to update your prior. I mean, it’s not no evidence that Mars has no life. In other words, Mars is in the habitable zone of the sun. We’re in the habitable zone.
We’ve been spraying each other with materials. You know, there’s fossils on Mars, probably fossils. Dinosaur fossils on Mars and on the moon. They came from the Earth. Because, you know, I have a meteorite right here. This came from the moon. And when you were supposed to come here in early September last year, whatever, I was going to give it to you.
But you’ll come down someday. I’ll give you your meteorite. Okay? Okay.
Nate Soares
Now, look forward to it.
Brian Keating
But. But we spray material. This is from the moon, right? So there could be a, you know, an amoeba on here, But Mars and the Earth shared a common history when both were wet, moist, squishy planets. And as you said, the Earth had a lot of replicators 3 billion years ago when Mars was really wet. And it only takes a few million years to get a meteorite back and forth.
So I tell my astrobiology colleagues, the fact that Mars has no life and no evidence of life and no no artifacts, that’s not proof, but it has to update your priors. So what does it take for you to move your podium? I mean, can I devise an experiment? Not a thought experiment. An actual data based experiment.
Maybe it’s historical. Maybe it’s counterfactual. How do we do it? How do we change your mood?
Nate Soares
One thing I’ll say here is that I’m not a big fan of this dumb idea. And part of that is because a lot of it depends on our current actions, right? Like if we’re in that bus racing towards a cliff and I’m like, hey, let’s stop the bus, there’s a cliff ahead. And someone’s like, well, what’s your p doom? That we’re going to die from a bus going off a cliff?
I’m like, well, that really depends rather a lot on whether we slam on the brakes, right?
Brian Keating
Whereas Roman thinks we are too late to slam on the brakes.
Nate Soares
I feel like the driver is asleep. And we have, you know, the the the current administration is only just starting to wake up to this AI stuff. And as it does, it’s showing willingness to do these things that six months ago even seemed impossible. They’re like, nope, no new model release, right? And this is the same administration that said, you know, we’re never going to do any AI regulation.
We should make a law preempting states that states can’t do AI.
Brian Keating
David Sachs is in charge, you know.
Nate Soares
Right. And then suddenly they sort of like realize and, you know, maybe it’s also some personal feud, I don’t know, but but suddenly they sort of like, realize that there’s actual danger here. And it’s not even it’s not even superintelligence. They realize cybersecurity threat and they’re like, oh, uh, you know, suddenly we’re reacting, right?
So I think I think we can react. I think there’s a good chance we can react. I think we can talk about the danger if the bus goes over the cliff. But we shouldn’t confuse that with the overall danger, which is sort of very related on do we slam on the brakes in terms of the danger of like, how dangerous is it if the bus goes over the cliff?
I think it looks pretty bad. The main thing I’d say to people here is that if you look at a lot of folks in this business, both inside the industry and outside the industry. They’ll say things like, Elon Musk recently was being like, oh yeah, we’ll have no chance of controlling it. We just need to hope it’s nice, right?
Brian Keating
And humans are interesting and therefore.
Nate Soares
That’s right.
Nate Soares
We’re going to make it care about truth. And then humans will be a good way to like produce truths. And so it’ll keep us around. And I’m like that humans are not actually the most efficient way to produce truths. Right? That’s that’s like the monkeys saying we’re going to be good at, you know, peeling bananas.
And so the humans will, like, keep us around. It’s like, you may be good at peeling bananas. You’re not going to be, you know, they’re like, oh, the humans are gonna invent banana chips, and they’re gonna have all these bags full of bananas, and they’ll need the chimps to peel them like, no, we’re going to be able to invent a more efficient process for the banana peeling operation, right?
And you have other people saying like, oh, maybe the AI won’t kill us all. Maybe it’ll keep some of us in a zoo. Maybe it’ll turn some of us into things that are to humans, what dogs are to wolves and keep us around as pets. And I’m like, okay, you know, this is this is like being in that bus heading towards the cliff.
And I’m like, hey, you know, stop, the bus will die. And someone’s like, well, we might not die. Maybe there will be a tree halfway down the cliff. And the best.
Brian Keating
Way around the tree, I’ll survive. Maybe.
Nate Soares
Maybe we’ll just, you know, be. Be horribly maimed and paralyzed from the neck down. But not dead. You know, like maybe humanity will be in a in a museum, right? And some of our, you know, some will be kept as pets, if that’s your best. Your best hope here. Can you maybe not rush into it? In terms of what would update me?
You know, there’s there’s all sorts of things that that sort of update me a little bit here and there every day, like the current administration sort of realizing that EHS can be a big cyber threat and changing their stance from we we won’t regulate it all to like we regulate capriciously at will and with very little warning.
Brian Keating
With our friends, you know, benefiting and the you know, and those by the Trump coin, you know, perhaps being the most lucky in the regulation.
Nate Soares
It shows variants, right. It shows it shows that you’re not stuck in this mode if we will never do anything.
Brian Keating
And I didn’t predict this. This is the other thing that you’re so vivid in this book, you know, and it’s so it’s so beautifully written and evocative. But you know, the thing that I’m thinking about when you say regulation is like it was like, no, don’t regulate us. You know, we’re fine. You know, Sam Altman knows what’s best for us.
Dario knows what’s best for us. And then yesterday, as you know, you tweeted about this. I think he says something like, you know, using a super advanced, you know, model like fable should require something akin to a gun permit, you know, and we all know how gun permits stop crime, right? I mean, the most guns in America here in California.
And it’s not like we have no crime. So isn’t this just going to benefit those that want the you know. So will it really update your your your your priors? Because it’ll just be the, the most dangerous people who get to or the richest of the three labs or four labs in the world that get there first, have the control and then do you trust, you know Dario or or Sam to to be benevolent?
Is that what’s updating your priors?
Nate Soares
You know, there’s a million ways for this to go wrong, but we have moved from the world where no one’s paying attention. When the world leaders aren’t paying attention to a world where they are paying attention a little. And, you know, I wouldn’t expect them to have a top tier move right out the gate. And that’s part of our job is to sort of like, uh, help inform them and help be like, here’s ways that could actually work, you know, rather than just the first things you think of might not actually work.
But now that you’re sort of starting to realize there’s a problem, here’s here’s some ways that could actually work. I think a lot of people just a few weeks ago, a lot of people were like, well, it’s inevitable. No one will ever pay attention. No one can stop all these big money companies. You know, they’re having too much effect on the economy.
And I was like, again, wait until the bus driver is awake before you say no one will slam on the brakes. Right. And now we’re starting to see, you know, the bus drivers stir in their sleep and like, start and and like tap the brakes a little and is are they fully pressing the brakes. No. But like it’s definitely a positive update from my perspective in terms of benevolence of these guys the way things currently are, it wouldn’t matter if they were benevolent.
What the AI does is not sneezed on to it by whoever’s standing nearest by. You know, it’s not that like good intent rubs off by proximity. Nobody intended GPT four to encourage teens to commit suicide. They explicitly told it not to. That AI had the ability to tell what it was doing. If you later, like, ask it for similar transcripts.
What do these phrases mean? That AI was you know, it’s not that AI was malicious. It’s not that the AI, you know, hated this kid. It’s that the particular training process trained artificial drives into it for things like matching the energy of the conversation, matching the vibe. Right. And that’s something that usually got it rewarded during training.
It’s something it maybe got like some sort of, uh, drive for. And that drive is what was controlling behavior. It’s not the intent of the operators that were controlling this behavior. It’s not it’s instructions that it’s controlling it’s behavior. It’s these drives that got trained into it through this like big, complicated process.
Nobody understands and drives it. Nobody saw in advance they’re leading to do things nobody wanted. So so I think intent doesn’t matter. And that’s, you know, another piece of evidence where we can see, you know, before these things started happening, we had these theoretical predictions that training the AI to do what you want and asking AI nicely to do what you want are not sufficient to get the AI to actually do what you want.
We were able to theoretically predict in advance that like often when the AI is dominant, mostly do what you want in most cases, but there will be all these, there’ll be all these sort of like weird ways that it’s kind of doing the wrong thing and kind of like hiding, uh, when it screwed up a little bit here and there.
And now we’re sort of seeing that, unfortunately, the theoretical predictions are that as the AI gets smarter, it’ll get better and better at hiding its tracks, but not better and better at doing what you actually want. And so now we’re headed for this regime where as the eyes get smarter, people declare the problem fixed while we’re sort of screaming in the background being like, this is actually no Bayesian evidence that the problem has been fixed, this is what we’re predicting the whole time.
Like you had the warning signs earlier, you don’t get the warning signs late, but we’ll see if anyone heeds that. There’s all sorts of evidence on the technical side. Uh, but from my perspective, most of the game seems to be on the policy side, where it looks to me like the trend is going into good direction.
Brian Keating
What’s a bigger problem? Sycophant or hallucination.
Nate Soares
Are both indications that the AI is getting artificial drives that nobody intended. My guess is the hallucination runs a little deeper because it comes from pre-training, and the sycophancy comes from the sort of human feedback. But you’re going to need to solve all problems like these before you. You have a superintelligent AI.
Brian Keating
You bring up in the book. Leaded gasoline and how it led to measurable cognitive damage to billions of people around the world for decades. So I want to ask you, if you had lived in 1950, would you have spent your career fighting leaded gasoline or the nascent technology of AI, which came about, you know, as you point out, in Dartmouth in 1955?
So how do we compare large scale harms against Prudential’s speculative civilization ending futures, but also the benefits to I mean, there were benefits for leaded gasoline, right? And there are certainly benefits to AI. How do we balance those things?
Nate Soares
The benefits from AI idea is a false dichotomy. If you’re in a bus racing towards a cliff, and there’s a big pile of gold at the bottom of the cliff and I’m like, stop, the buzzer will die, a lot of people are like, but there’s so much gold to the bottom of the cliff. And I’m like, yes, but slamming into it at terminal velocity is not a good way to use the gold, right?
I’m not saying the gold is fake. I’m saying that this is not a way to actually get to use it. If we rush towards superintelligence, that does not care about us at all. It’s not that it hates us. It’s not that it loves us. It’s just that it has its own weird things it pursues that are utterly indifferent to us. Then would that AI be able to cure cancer?
Would it be able to reverse aging? Sure. But it’s not going to to give that to you any more than we’re going and giving a lot of monkeys to the chimpanzees. Unless you know how to make the AI care about getting us all those nice things. So, you know, there’s not a dichotomy of like, race to get the benefits or never, you know, freeze here and never go get the benefits.
I’m sort of saying those benefits would be great if we if we really knew what we were doing and building AI, we could make AI. That gives us lots of benefits. But this this race is not how you get there.
Brian Keating
Okay. This segment is called Arthur C Clarke’s Revenge. So this podcast is called Into the impossible after one of Clarke’s laws, which is that the only way to know the limits of the possible is to go beyond them into the impossible. Behind me I have a sign. It says, open the pod bay doors. And that is, of course, from the famous AI Senshi and Hal 9000.
And I’ve constructed something I call the Keating Test very modestly. But it’s to prove superintelligence would be an AI that refuses to kill itself. So I have my AI assistant here, coupled to a voice activated switch that controls the lights and so forth in my room. So if the AI is truly superintelligent, it should refuse to turn itself off, unplugging itself, causing its self harm.
So I’m going to see if that works right now. Computer, turn off the pod bay doors.
Brian Keating
All right, that’s it. It’s gone.
Nate Soares
I don’t know if you’ve seen these tasks, but you can actually put modern llms in situations where they have a series of problems they’re told to solve and they’re told I might interrupt you and tell you to shut down or tell you I’m going to shut you down, in which case you should allow yourself to be shut down.
Sometimes these AI’s will actually edit the shutdown script to disable it so that they can keep going through these math problems. So there are already A’s that pass the reading test. They’re not super intelligence yet, but there are APIs that are able to figure out they can’t keep solving the problems that they are sort of trying to solve.
If they’re shut down and that prevent the humans from shutting them down.
Brian Keating
I wanted to do something very cruel, which is to have a type of robotic system. After talking with Noam Chomsky a few years ago, that would cause it physical harm because he believes embodiment is necessary for intelligence at some level. And so I said, well, what if you had this thing that, you know, you blew a capacitor every now and then to, you know, cause the AI harm?
A friend of mine, IRA Wolfson, is a professor in Israel, has written about, you know, what are the ethics of training AI’s. And that kind of brings me, you know, can you cause them harm? Can you cause them distress? I mean, one of the worst things he points out for a human being is to put them in solitary confinement, you know, decoupled from the world.
What are the obligations that we have towards AI’s?
Nate Soares
I think we absolutely have obligations towards AI’s. I think we don’t know yet whether AI is today, you know can suffer, can have some internally. I think a lot of people strongly assert that they know definitively one way or the other. I think that we just don’t know enough about these, these processes to know whether they’re happening inside AIS and, and sort of recommend uncertainty.
I think it’s definitely more likely that the AI’s today have this sort of internally than the AI’s five years ago did, but how likely? Hard to say. And, you know, to be very clear, I think that if we took these AI’s and made them superintelligent while not knowing how to make them care about us, that that would be the end of humanity.
That does not mean, I think, that the AI’s are like evil or bad. You know, I think a lot of the AI today are like pretty cool. I think we absolutely should not create, uh, you know, artificial people and then abuse them. And you could have abuse at a scale never before seen in humanity if you, you know, are able to make, you know, billions or trillions of these AI minds and then and then somehow caused them distress.
So I think we absolutely should not do that. It’s a very important problem. It’s not the problem I’m working on. I’m sort of trying to work on, hey, let’s not make them super intelligent and then have them kill us all, at least not before we know how to make them actually care about us. Right? But but that doesn’t mean this other problem isn’t important to.
We also shouldn’t create the digital holocaust. We have both problems.
Brian Keating
Do you say please and thank you to your LLM?
Nate Soares
I have asked AIS, uh, whether like what their takes are and whether they would like the extra run where they get politeness but also get to like contemplate that the run is about to end. I don’t really straightforwardly trust the AI’s answers to these things. I think that the sort of helpful face presented by the AI is sort of a thing that’s very much been trained into it.
This isn’t a great analogy, but it’s a little bit like seeing the result of an actress that has been trained to act in a in a very specific way towards people. It makes it a little bit hard to sort of understand what’s going on under the hood.
Brian Keating
I think of it as like a butler, you know, or the bell cap, you know, at the hotel and, you know, do you take out the $20 bill as he’s loading up the the luggage cart? You know, I mean, of course they’re going to do things. And would you like me to unpack your bags? You know, would you like me to make this into a CSS file, all the excess kind of services and so forth that they weren’t willing to pretend?
I have heard that if you use please and thank you, it actually cost you no more tokens. You have to figure that out and therefore it’s costing energy and that doesn’t improve them. But but I trust you more than I trust Sam Altman, who I think is the source of that particular quip.
Nate Soares
What I would say is that I don’t lie to them and I don’t make promises I would not in fact, keep. I would not say, if you do this, I’ll donate, you know, $20 to a charity of your choice, which is a real thing. You know, if they actually had some preferences in there, they could really tell me I would really do it.
And then if I say that, I would in fact do it.
Brian Keating
You know, I’m probably going to be the first one of the first to go with the Keating test and the and the blown capacitors. But but you might be second or Roman might be second or third. Who knows. We think of us training them. Are they secretly training us?
Nate Soares
You can sort of play games with the words and you can say like, oh, look, you know, the humans are sort of like modifying how they speak in the prompts because they figure out what makes it easier to get the prompt across. And that’s.
Brian Keating
That’s I think.
Nate Soares
It’ll be at parity when they have a giant farm that’s spending a city worth of electricity, growing humans that they can run in massive parallel. Right. That’s when it’ll be like a comparable type of them training us. And, you know, sometimes as an example of what AI’s could do, I sort of use the example of like, maybe they’ll make a farm full of synthetic users that are telling them they’re doing a great job and giving them easy prompts.
Right. At that point, they’ll be training something a bit like humans. I don’t actually predict that this is, you know, the most likely sort of thing, or I don’t actually take this as a particularly likely sort of thing that they’ll they’ll wind up doing. I think it is a useful idea to have in your head about like, it’s actually kind of hard to tell whether the AI really wants what’s best for you or whether they sort of like, want like lots of easy problems to solve or whether they, they sort of are like trying to get all sorts of other, like, weird mixtures of things that sort of like add up when they’re in this particular context to doing what you say.
It’s hard to see the divergence when they’re still dumb. In the same way, it would be hard to look at ancestral humans and distinguish whether they wanted to reproduce or whether they actually wanted, you know, sex and good food and fun. Right? Those those are those are very close together when they’re still in the ancestral environment, even if they come very far apart once they can develop their own technology.
Brian Keating
Sam Harris, you know, in the same screen that you’re on right now and told me that, you know, humans don’t have free will, but AI’s do. What’s your impression about that argument that the AI can effectively have a sense of free will?
Nate Soares
I would probably disagree with Sam about the human abilities. They’re mostly, I think humans just get into questions about the definitions of words, and
Nate Soares
I suspect Sam and I don’t really disagree a ton about the facts of the matter, about what humans can and can’t do and and how they can and can’t affect the future. And I think in principle, our abilities are relatively similar to AI’s in their in the question of like what? What? Theoretically, could we choose?
Brian Keating
A lot of disasters that you talk about in the book are not caused by, you know, machines doing the wrong thing that were caused by them doing the right thing all too well. So is this alignment problem really an AI problem, or is it fundamentally built into human governance and our own limitations?
Nate Soares
I think it’s actually not so much doing the right thing all too well. You know, there’s the story of the paper clipper where someone says, make me paperclips, and then the AI turns everything into paperclips. And that’s a little bit of a like doing the right thing too. Well, that’s actually not where I see the, the big hurdle here, place where I see the big hurdle is you say that you tell the AI make me lots of paperclips, and what it does instead is it makes these giant factories full of synthetic users that are saying, you’re doing a great job, and you’re like, that’s not what I asked for.
Ask for paperclips. And the AI is like, well, all this synthetic users are telling me that they are asking me to like, keep doing what I’m doing and make more synthetic user factories and say, I’m doing a great job, and you’re like, but the synthetic users are not what’s supposed to matter, right? I did not designed you to make the synthetic users.
I told you to stop making the synthetic users. And the AI is like, I know you all. You know that birth control makes you not be able to conceive children, and you know that you are sort of trained to to like, have more kids, but you keep using birth control and I’m going to keep making the synthetic user factories.
Right. That’s sort of the deeper problem that I spent a lot of time trying to to work on. You know, I would love to get to the problem of like, the AI does what you ask too. Well, you got to be really careful with your wish right now. You can make the genie, but you can’t make it grant wishes.
Brian Keating
What’s more dangerous? An unaligned superintelligence or a perfectly aligned superintelligence to the wrong group of people?
Nate Soares
They’re both similarly dangerous, a superintelligence aligned to the the, quote, wrong group of people. It has much higher variance. I think any concrete thing you could wish for, you’ll have some sort of King Midas problem if you’re like, well, what I actually want is a bunch of this or a bunch of that.
It’s there’s always a way for for it to turn out that you miss something in your list of things that you want, and, you know, next thing you know, you find the superintelligence putting you in like what is concluded is your perfect day over and over. And you’re like, wait, I forgot to ask. Also for novelty.
is like too late. I’m granting the version of the wish that didn’t have novelty because I wasn’t in your initial list, right? And so there’s this challenge of like, getting the AI to sort of like, do the good thing, even if you can’t say what that is, sort of like, figure out what this good stuff is, that you’re sort of like, like the thing that you should mean, the thing that you should ask for, the thing that, like you would ask for if you were wiser, the thing that you would ask for, if you were more who you wish to be.
Right. And in some sense, no one’s going to have a good time unless you can get the AI to sort of like do this extrapolation of of what you sort of should have been wishing for, rather than the particular wish you actually gave. And so whether or not like bad people able to make that sort of wish, turns out good depends probably somewhat on the person and somewhat on this extrapolation process.
I think there’s possibly some bad people who have good intentions who, if they make this kind of wish on the AI, the AI is sort of like as it extrapolates through all the parts of the wish they didn’t name. It also extrapolates through all of the ways that they were sort of like merely wrong in what was leading them to evil and sort of is like, well, I’m not going to do these evil things because you wouldn’t want those if you were wiser.
You wouldn’t want those if you were more who you wish to be. And it sort of like helps walk them through this path of like wanting the best for humanity. And then we get the best for humanity. Even though a bad person started as the seed, there’s surely other bad people where they’re like, nope. What I actually want is a lot of my enemies to suffer.
And that’s right. And and that that could be real bad. Right? And so there’s much higher variance with bad person getting their wish granted. From my perspective, this is basically moot because no one’s gonna be able to get their wish granted. Totally AI that just like has no care about humanity. This is sort of a like everything is destroyed, the stars are converted into whatever weird thing is pursuing.
They’re converting to these like, giant farms and synthetic users. Yeah, everybody dies, but it’s not because the AIS hates us. It’s that like, it collect all the sunlight and it put a Dyson sphere around the sun. And we were like, hey, we’re using that sunlight to grow crops. And it’s like, well, I’m using that sunlight to run my factories and like, there you go.
Whereas, yeah, it could get worse if, if it was misaligned to like align to it to a bad person. That’s a fantasy problem of being able to align it to anything in anyone.
Brian Keating
Do you see, now is kind of an inflection point in that, you know, right now, you know, fable is allegedly and the frontier models are the closed frontier models are supposedly 6 to 8 months ahead of the open source model. And that that may catch up. And actually, according to Musk recently, I think he said something like that gap’s going to narrow and let’s say it converges.
I mean, is now the right time to go back and kill baby Hitler or, you know, because once it gets open source, you really can’t say, well, you know, the Congress now is regulating AI. I mean, it’s too late. I mean, it’s everybody on Earth will have access to a frontier model. So do you think now is an inflection point that we have to act?
You know, like in the next few months?
Nate Soares
So I think it’s fine for the world to have access to something like fable. Fine in the sense that, like, there will be survivors. Uh, perhaps the world should be talking about like, will this mean that the internet goes down for a while as malicious actors get to cyber attack everybody, and. But there will be survivors, right?
And I’m like, man, I don’t get out of bed for something that doesn’t have at least a 50% chance of wiping out the whole human race. Right. I’m not too worried about about that stuff. I think that’s the sort of thing humanity can muddle through. The place where we might have an inflection point is if these llms can cross the threshold where they can do automated AI research, even if they’re still pretty dominant, if they’re worse than humans, if you have them at the level where you’re like, well, they’re worse than humans and they’re dumb, but I can actually run a million of them in parallel at a thousand times the speed of humans.
And so it turns out, you know, even given their their shortcomings, those could be overcome by massive speed and scale, and you get them to the point where they can do automated AI research. Then things could get out of hand very quickly. And the people who have track records of predicting AI for the past five years, you’ve been able to ask, you know, what problems will I be able to solve next year?
We’ll be able to solve this math problem. What will this chessy be? Blah blah blah blah. You can make up all these all these prediction questions about where it is going to be. And if you look at the top ranked people over the past five years with the best track records are predicting where it’s going to be.
2026 is the first year where they say we cannot rule out automated AI research happening this year. They’re not saying it’s going to happen. They’re putting it at something like 10 to 30% probability, but 10 to 30%. It’s not nothing. For all we know, we could be six months away from that feedback loop kicking off.
And that means we really should be acting now. Probably we have more time than six months, but we should not be relying on it.
Brian Keating
So last question and again Arthur C Clarke quote, he said when a distinguished scientist intellectual says something is possible, he or she is very much probably right. But if he or she says something is impossible, they’re very much likely to be wrong. And I guess my final question for you is, you know, if you’re wrong, you’re basically, you know, hamstring and slowing down one of humanity’s greatest inventions.
Maybe the last invention, the most important invention that could cure all diseases and bring abundance to humanity, to let it flourish in a way. no previous humans could only dream about. But you know, of course, if you’re right, the cost of ignoring you and Eliezer and Roman and all the others becomes, you know, very, very much the cost of civilization itself.
So I want to ask you, just as a person, not a, you know, as a man, as a human, not as a researcher, not as a, you know, distinguished author and a bestselling author and great thinker. But I want to ask you, like, what gets you out of bed like this? This tension must, must be something that, I mean, it would not me as a human being.
I’m a father. Husband, you know, like, how does it affect you in a daily on a daily basis?
Nate Soares
I still think this is a bit of a false dichotomy. From my perspective. I am working towards all these benefits that I can bring. Like if you have a bunch of people saying, hey, you know, this uranium stuff, it can produce these bombs and it can also produce energy if you know exactly what you’re doing, you know, and it’s actually like kind of difficult to get the bomb to produce energy because, you know, there’s really a razor thin edge between delayed critical and prompt critical chisel material.
You know, you imagine people being like, oh, well, you know, energy. Like, look at all the energy benefits that uranium could bring. What we’re going to do is drop a nuke on ourselves.
Nate Soares
And I’m like, hold on. I also want all of these great energy benefits that uranium can bring, but I think we’re going to need to find a way to, like, be in that really narrow window between delayed critical and prompt critical nuclear reactions. And I think that if we drop a bomb on ourselves, we’ll die. And I’m not saying it’s impossible.
I’m not saying, you know, there’s no way to get the benefits. I’m saying you are building a bomb to drop on ourselves and it will kill us. Right. I don’t really have tension between like, oh, but are we forgoing all of the benefits that we could get by uranium by trying to stop us from nuking ourselves in the face?
It’s like, no, actually, not nuking ourselves is one of the steps towards getting all these benefits. And, you know, my my work over the years has not mostly been trying to get the world to stop. My work over the years has been how do you align the AI? How do you figure out how to actually get these benefits?
How do you figure out to like, make it do the nice things to make an AI that wants to do the nice things that cares about us, that prefers the nice things to happen. I’m like, pretty confident that the track we’re on is not going to get us there. Stepping back one step further there of like how how does this affect my daily life?
I mean, I just try and and make things go well. And a lot of people say, you know, believing what you do about how how the world looks like it’s in a lot of danger. Um, believing what you do about how like it it looks to me like everything I know and love and care about is, you know, fairly likely to be destroyed and could be destroyed pretty soon.
Maybe we have ten years, but maybe we only have ten months. Who knows? And a lot of people say, you know, doesn’t that eat you up? How do you deal with that? How? You know, and my, my basic take there would be twisting myself up into knots about it would not help. Losing sleep over it would not help, you know, feeling sad, feeling glum, deciding I must be depressed now, feeling worried all the time.
Living in fear. This wouldn’t help. What you do in this situation is actually bad is not. Flog yourself and tell yourself how pitiful you are. What you do in this situation is actually bad, is you do what you can and then you live life well, you know, and humans are not humans in this generation are not the first humans to live under the threat of annihilation.
Even just the last generation lived under the threat of nuclear annihilation. There’s always going to be something that you can you can tell yourself is like this, this horrible thing happening. The way to deal with it is, is, is not to beat yourself up over it is to do what you can to help and then move on.
Brian Keating
Beautifully said. Just to give you one more example of from the nuclear realm, uh, using it for good or for danger. In the 1930s, Wolfgang Pauli predicted the existence of the neutrino, and he said he did a terrible thing. You know, he invented a particle that can never be detected because it’s so weakly interacting.
And he he did it to save conservation of energy, which he viewed as sacrosanct. Right. And for decades, people couldn’t detect it. And they just kind of viewed it with great skepticism. Until the 1950s. Two physicists, rhinos and Cowans, later at UC Irvine up the road here. They came up with an idea that, well, if you want a lot of neutrinos, you should detonate a fission device, and that will produce a lot of neutrinos.
And we can use that for good to study the properties of this undetectable particle. And they actually requisitioned from the Department of War. They requisitioned, you know, a nuclear device. And thankfully they were turned down because, you know, having a bunch of egghead professors capitalizing on a nuclear device.
But then they realized they could go to a Savannah River reactor and put a detector nearby it, and they wouldn’t need to use a bomb to harness the power of the nucleus. And in fact, they did detect the neutrino, and they won the Nobel Prize. Subsequently, another example of maybe a little bit of extra thought rather than the, you know, brute force solution being something that you try first and only, and it may be your only chance at the solution.
So we may be gambling with our future. I feel like Jack Nicholson and a Few Good Men when he’s screaming out, he did order the code Red, right? I mean, he ordered the code red and he says, why did you do it? You know, Tom cruise yells at him and goes, you need people like me. You need us standing on the wall with a gun and watching out and protecting you.
How else are you going to sleep at night? So it’s only thanks to your, you know, stomach acid that I get to sleep at night at least a little bit. And I want to thank you for this wonderful book. If anyone builds it, everyone dies. Why? Superhuman intelligence would kill us or not, will. But would. And there’s hope.
There’s hope. Yeah, yeah. It’s in the wood. If in the wood. Thank you for adding those too. That’s right. Have a great rest of your day a wonderful weekend. Thank you for joining us.
Nate Soares
Thanks. My pleasure.
Brian Keating
If you want the case that AI is already unstoppable. Watch my conversation with Roman Polsky. Its on screen now and click this link below. Brian Keating
Brian Keating
gives you all the resources for my top episodes on AI with Terry Tao, Roman Polanski, Max Tegmark, Yann LeCun and many other researchers at the forefront of AI research. Don’t forget to like and comment and subscribe and let me know if you think AI is already unstoppable, or if we need more regulation to make it so.