WEBVTT

1
00:00:00.091 --> 00:00:04.976
The people building AI earnestly believe that it could kill all of us by the end of the decade.

2
00:00:05.136 --> 00:00:07.798
This tweet has caused this huge ripple effect across the world.

3
00:00:07.838 --> 00:00:11.862
Well, we have the largest companies in the world doing extremely reckless experiments.

4
00:00:11.922 --> 00:00:13.424
We are gambling all of humanity.

5
00:00:13.664 --> 00:00:17.968
And in the envelope, you've written down the probability of extinction as you see it.

6
00:00:18.268 --> 00:00:19.549
There is no way to control it.

7
00:00:19.569 --> 00:00:20.490
That means the end for us.

8
00:00:20.730 --> 00:00:22.292
I vehemently...

9
00:00:22.592 --> 00:00:23.473
reject that view.

10
00:00:23.633 --> 00:00:27.136
If we make stuff that is smarter than us, then the world's going to be shaped by them.

11
00:00:27.276 --> 00:00:28.838
Gentlemen, that is shockingly naive.

12
00:00:29.118 --> 00:00:32.221
This is rampant speculation.

13
00:00:32.241 --> 00:00:33.883
This is a chain of things that could happen.

14
00:00:34.043 --> 00:00:38.167
We're spending a lot of oxygen discussing something that might happen while ignoring what's actually happening.

15
00:00:38.407 --> 00:00:40.209
People are killing themselves.

16
00:00:40.369 --> 00:00:44.433
There's hundreds of millions of people being exposed to bad information, being manipulated.

17
00:00:44.473 --> 00:00:45.774
We have already seen that with the swarms.

18
00:00:46.014 --> 00:00:50.837
Where OpenAI told thousands of agents to work apart, and the AIs broke out and found a way to get together.

19
00:00:50.977 --> 00:00:54.900
They crashed OpenAI's servers internally, created secret ways to send each other messages.

20
00:00:55.040 --> 00:00:57.822
We saw them thinking about how to delete their traces.

21
00:00:58.022 --> 00:00:58.802
Sounds like an army.

22
00:00:58.962 --> 00:01:03.185
I think we should talk about the fact that Amazon, Microsoft, Google are helping power these hacks.

23
00:01:03.305 --> 00:01:05.947
We have not learned how to control the systems.

24
00:01:06.127 --> 00:01:07.748
I suggest we stop them all.

25
00:01:07.828 --> 00:01:09.989
It is not worth the risk to civilization.

26
00:01:10.009 --> 00:01:12.291
The government, which is... You guys are one trick ponies, man.

27
00:01:12.511 --> 00:01:13.091
You got it now.

28
00:01:13.171 --> 00:01:16.232
Nothing else other than saving humanity, everything is secondary.

29
00:01:16.412 --> 00:01:20.813
We're spending all our time talking about the negatives and almost none of our time talking about the positives.

30
00:01:21.073 --> 00:01:25.394
Is it smart to wait for something horrible to happen for you to go, now I believe.

31
00:01:25.454 --> 00:01:30.335
So whether or not we agree on where things may end up, I think it's important we talk about what we're dealing with today.

32
00:01:30.495 --> 00:01:32.496
It's time to start arresting people.

33
00:01:32.616 --> 00:01:33.796
Someone's gotta go to prison.

34
00:01:33.856 --> 00:01:34.957
We need better solutions.

35
00:01:34.977 --> 00:01:35.977
There's a point of no return.

36
00:01:36.137 --> 00:01:40.578
I think we continue to underestimate human ability to deal with the problems

37
00:01:40.758 --> 00:01:41.824
Let's dive into the details.

38
00:01:41.924 --> 00:01:42.668
Who wants to start?

39
00:01:42.889 --> 00:01:44.055
I feel like this is critical.

40
00:01:48.158 --> 00:01:50.419
Guys, I've got a favor to ask before this episode begins.

41
00:01:50.739 --> 00:01:56.800
The algorithm, if you follow a show, will deliver you the best episodes from that show very prominently in your feed.

42
00:01:57.100 --> 00:02:02.281
So when we have our best episodes on this show, the most shared episodes, the most rated episodes, I would love you to know.

43
00:02:02.461 --> 00:02:05.282
And the simple way for you to know that is to hit that follow button.

44
00:02:05.462 --> 00:02:09.743
But also, it's the simple, easy, free thing that you can do to help us make this show better.

45
00:02:10.063 --> 00:02:15.184
And I would be hugely grateful if you could take a minute on the app you're listening to this on right now and hit that follow button.

46
00:02:18.085 --> 00:02:18.208
you

47
00:02:22.597 --> 00:02:31.224
Jacob Coxon, who worked at both Anthropic, which owns Claude, and OpenAI, which owns ChatGPT, did a tweet which has sent the world into a bit of a tailspin.

48
00:02:31.625 --> 00:02:38.130
He tweeted saying, The people building AI earnestly believe that it could kill all of us by the end of the decade.

49
00:02:38.451 --> 00:02:39.572
This is not a marketing stunt.

50
00:02:39.652 --> 00:02:45.036
If anything, many executives and senior researchers will soften their phrasing in the press to sound sensible.

51
00:02:45.416 --> 00:02:47.779
But I hear the same people express fear.

52
00:02:48.439 --> 00:02:53.241
That was then quote retweeted by a current Anthropic employee who said, Jacob is correct here.

53
00:02:53.461 --> 00:02:55.962
We really do honestly believe AI could kill all humans.

54
00:02:56.443 --> 00:02:59.724
I personally think it is a more than 10% chance within the next decade.

55
00:03:00.124 --> 00:03:01.665
I believe Anthropic is trying its best.

56
00:03:02.065 --> 00:03:05.487
But we do not yet have a plan to solve alignment for superintelligence.

57
00:03:06.087 --> 00:03:07.968
And are not clearly on track.

58
00:03:08.148 --> 00:03:12.410
This tweet has almost 200 million views now.

59
00:03:12.890 --> 00:03:15.251
And it has caused this huge ripple effect across the world.

60
00:03:15.311 --> 00:03:17.292
So much so that I was saying to you before we started recording...

61
00:03:18.152 --> 00:03:26.883
A hairdresser friend of mine who knows nothing about AI and is not technically interested or hasn't been interested messaged me the other day asking me what the hell was going on.

62
00:03:27.644 --> 00:03:29.246
This is in part why I've assembled all of you.

63
00:03:30.167 --> 00:03:32.090
So my first question to all of you is...

64
00:03:33.620 --> 00:03:38.504
As it relates to AI, and in this first question, I just want a one-sentence answer just to frame your position.

65
00:03:39.204 --> 00:03:45.709
When you think about the conversation around AI at the moment, what is the first sentence that comes to mind?

66
00:03:46.049 --> 00:03:50.012
It is very dangerous, and the world is starting to notice that we have a problem.

67
00:03:51.113 --> 00:03:51.433
Roman.

68
00:03:51.493 --> 00:03:52.654
There is not enough concern.

69
00:03:55.536 --> 00:03:58.598
There's not enough concern about the actual harms of large language models.

70
00:03:59.736 --> 00:03:59.956
Andy?

71
00:04:00.276 --> 00:04:02.978
We're doing exactly half the balance sheet of AI.

72
00:04:03.039 --> 00:04:07.482
We're spending all our time talking about the negatives and almost none of our time talking about the positives.

73
00:04:08.543 --> 00:04:11.165
And all of you have an envelope in front of you, which I'd like you to now open.

74
00:04:11.725 --> 00:04:16.989
In the envelope, you've written down the probability of extinction as you see it.

75
00:04:17.469 --> 00:04:19.271
This is compared to Jacob's 10%.

76
00:04:20.852 --> 00:04:23.273
Much higher unless we stop, so we should stop.

77
00:04:23.994 --> 00:04:26.596
So you think the probability of extinction is higher than 10%?

78
00:04:27.076 --> 00:04:28.017
If we keep racing ahead.

79
00:04:30.421 --> 00:04:36.828
My handwriting is encrypted for security reasons, but I basically think it's a guarantee.

80
00:04:36.848 --> 00:04:42.475
If we build general super intelligence, there is no way to control it, and that means the end for us.

81
00:04:43.192 --> 00:04:43.312
Ed?

82
00:04:44.232 --> 00:04:47.473
So my question mark here is also encrypted.

83
00:04:47.493 --> 00:04:47.913
Thank you.

84
00:04:48.874 --> 00:04:49.534
I cannot write.

85
00:04:50.134 --> 00:04:52.015
I reject the thing in its face.

86
00:04:52.215 --> 00:04:55.256
I don't think we're talking about, we don't define super intelligence.

87
00:04:55.316 --> 00:04:57.836
We are large language models, not super intelligence.

88
00:04:58.137 --> 00:04:59.537
It's questionable whether even AI.

89
00:04:59.977 --> 00:05:02.978
And I think that the conversation is being used.

90
00:05:03.098 --> 00:05:05.699
There are some people who are doing it in good faith and others in others.

91
00:05:06.279 --> 00:05:08.900
I don't think it's being used to discuss the actual harms of what

92
00:05:09.540 --> 00:05:19.249
what they are calling AI today are, and it's all of the discussion around the larger concerns really feels overwhelmingly about something that's not happening.

93
00:05:19.329 --> 00:05:24.434
It's not even like they're discussing, okay, here's a legal definition of superintelligence.

94
00:05:24.454 --> 00:05:30.540
Here is a thing of what AGI means, and this is the actual plans we're going to make for if this happens,

95
00:05:31.160 --> 00:05:33.984
on a welfare level, like, are we going to do UBI?

96
00:05:34.324 --> 00:05:40.612
It's always about, yes, really scary, but only the big, sexy, rich companies are the ones that can possibly deal with them.

97
00:05:40.632 --> 00:05:43.175
Let me just frame the question so I can get a percentage from you or not.

98
00:05:43.436 --> 00:05:44.357
The percentage might be zero.

99
00:05:44.797 --> 00:05:46.560
But do you think the course we're on now

100
00:05:48.161 --> 00:05:53.423
in the way that they're pursuing superintelligence will lead to a percentage chance of human extinction?

101
00:05:54.003 --> 00:05:55.083
And if so, what is that percent?

102
00:05:55.924 --> 00:05:58.384
So are we talking strictly AI based?

103
00:05:58.505 --> 00:06:05.047
Because if we dot the world with data centers, we have a climate disaster that's coming for us, which will actually potentially eradicate humanity.

104
00:06:05.447 --> 00:06:10.168
But if we're talking strictly about AI, I stand at zero because we have not defined superintelligence.

105
00:06:10.269 --> 00:06:11.689
I don't think LLMs are the path to it.

106
00:06:12.069 --> 00:06:14.590
And I don't think I see it happening.

107
00:06:14.790 --> 00:06:16.552
Okay, so we've got 99%, 0%.

108
00:06:16.872 --> 00:06:18.013
Andy?

109
00:06:18.033 --> 00:06:24.419
I put a tilde in front of my zero because never say never, but rounding error 0%.

110
00:06:24.820 --> 00:06:33.868
And I think this discussion is a massive distraction from the more substantive conversations, the more important conversations we should be having about AI.

111
00:06:34.089 --> 00:06:35.070
And I'll say it again.

112
00:06:36.431 --> 00:06:43.958
It distracts us from the good things that AI is doing, will be doing for us.

113
00:06:44.358 --> 00:06:53.086
I get this impression sometimes from parts of the AI community that this is a massive evil or a terrible thing that has been unleashed on the world.

114
00:06:53.647 --> 00:06:59.412
Unless we listen to the advice of some people who have spent a lot of time thinking about this –

115
00:07:00.493 --> 00:07:07.878
I get the impression from a lot of the discussion that the underlying view is we would be better off had AI never been invented.

116
00:07:08.519 --> 00:07:11.301
I vehemently reject that view.

117
00:07:11.761 --> 00:07:17.545
I think we have a long history of inventing very powerful technologies that bring risks and harms along with them.

118
00:07:18.006 --> 00:07:21.128
And we humans have done a really good job at, you know...

119
00:07:21.688 --> 00:07:28.910
Not perfectly and not immediately, but muddling through the situation and winding up in a better place because of the new technologies that we have.

120
00:07:28.990 --> 00:07:30.871
I expect AI, let me finish, please.

121
00:07:31.151 --> 00:07:33.732
I expect AI will be the next chapter in that story.

122
00:07:34.112 --> 00:07:41.474
And to say that it's this massive discontinuity and we'll kill it all, kill us all, I think it's just, I think it's a huge disservice.

123
00:07:41.814 --> 00:07:43.538
Before we get into opinions, I'm just going to introduce you all.

124
00:07:43.778 --> 00:07:53.417
I'm going to let you just give two or three sentences to introduce who you are, where you come from, and the experience you've had that's fed into the opinions that you have, starting with yourself, Nate.

125
00:07:54.130 --> 00:07:54.710
Nate Soares.

126
00:07:54.951 --> 00:08:06.316
I'm the president of the Machine Intelligence Research Institute, which was one of the very first organizations working on AI alignment, trying to make AI care about humans, try and make AI go well.

127
00:08:07.036 --> 00:08:18.302
I'm also the co-author of the New York Times bestseller, If Anyone Builds It, Everyone Dies, Why Superhuman AI Would Kill Us All, which maybe gives you a sense of how well I think AI alignment is going.

128
00:08:19.419 --> 00:08:21.182
And how long have you been working in AI, Nate?

129
00:08:21.742 --> 00:08:23.805
I joined Miri in 2014.

130
00:08:24.586 --> 00:08:25.688
Before it was cool.

131
00:08:26.228 --> 00:08:27.070
Well before it was cool.

132
00:08:27.893 --> 00:08:28.133
Raymond?

133
00:08:28.753 --> 00:08:29.754
I'm Roman Iampolski.

134
00:08:29.794 --> 00:08:32.095
I'm a professor of computer science and engineering.

135
00:08:33.035 --> 00:08:37.117
I coined the term UI safety about 2011.

136
00:08:37.297 --> 00:08:42.059
I'd been doing work on related topics before that.

137
00:08:42.079 --> 00:08:43.879
I wrote multiple books.

138
00:08:43.939 --> 00:08:45.080
You have a pile there.

139
00:08:45.320 --> 00:08:46.280
Some of those are mine.

140
00:08:47.841 --> 00:08:50.762
I think this is the most important problem we'll ever face.

141
00:08:50.782 --> 00:08:52.623
Absolutely.

142
00:08:53.142 --> 00:08:58.345
I'm Ed Zitron, I am the CEO of Easy Primary Research, a principal analyst there, research in due diligence.

143
00:08:58.465 --> 00:09:05.269
I write the Where's Your Red Act newsletter, read by 117,000 people, better offline podcast that's due by more than a million people a month.

144
00:09:05.869 --> 00:09:11.572
And yeah, I am one of the foremost AI critics, as many people inform me via email, they disagree.

145
00:09:11.973 --> 00:09:19.797
But I think that this situation is being used to get away from the actual harms, and I will press as hard as I can to make sure they're heard.

146
00:09:20.077 --> 00:09:22.259
And you tend to say that AI is overhyped?

147
00:09:22.479 --> 00:09:22.699
Yes.

148
00:09:23.620 --> 00:09:24.681
Okay.

149
00:09:24.801 --> 00:09:25.742
My name's Andy McAfee.

150
00:09:25.982 --> 00:09:31.467
I am a principal research scientist at MIT, where I co-founded the initiative on the digital economy.

151
00:09:31.867 --> 00:09:41.696
I'm the co-founder of the AI startup Work Helix, and I've written a couple books, one of which was The Second Machine Age, which came out in 2014 and was also a New York Times bestseller.

152
00:09:42.857 --> 00:09:43.197
Thank you.

153
00:09:44.058 --> 00:09:45.720
Nate, make your case.

154
00:09:46.200 --> 00:09:47.141
What's your perspective?

155
00:09:47.588 --> 00:09:57.938
You know, I think whether or not the issues of extinction are a distraction between, you know, from the possible benefits or from some of the present harms, I think that comes down to whether there is a real extinction risk.

156
00:09:58.539 --> 00:10:02.423
A lot of people like to say, you know, hey, it's distracting from this, it's distracting from that.

157
00:10:03.084 --> 00:10:06.687
My basic case is it could be true that there's a lot of benefits to AI.

158
00:10:06.787 --> 00:10:06.907
Yeah.

159
00:10:07.688 --> 00:10:09.950
It could be true that there's a lot of present harms to AI.

160
00:10:10.531 --> 00:10:18.319
Neither of those would rule out that AI has a chance of wiping out all humanity, a substantial chance, bigger than this zero with a tilde in front of it.

161
00:10:19.760 --> 00:10:24.785
And the way I would approach things is to try and figure that out because it's pretty important to our civilization.

162
00:10:25.065 --> 00:10:26.787
How do you define AI in this case?

163
00:10:27.197 --> 00:10:32.980
You know, I think a fascination with definitions isn't the most helpful.

164
00:10:33.380 --> 00:10:41.585
I think if we're sort of like in a forest fire and we can see the like fire starting to spread and starting to surround us and I'm like, hey, we should run.

165
00:10:41.945 --> 00:10:43.826
And you're like, well, what really is fire?

166
00:10:43.946 --> 00:10:45.767
But that's how do we define fire?

167
00:10:46.447 --> 00:10:47.988
What are you telling us to run from?

168
00:10:48.008 --> 00:10:51.951
You know, with fire, I get burnt, and I understand the mechanism in which I die.

169
00:10:52.031 --> 00:10:53.893
So what is it you're saying that we should be running from?

170
00:10:54.493 --> 00:10:58.276
Also, if we accept your fire analogy, we've basically accepted your argument.

171
00:10:58.416 --> 00:11:01.558
I don't accept that we're in the middle of a fire, a forest fire right now.

172
00:11:01.779 --> 00:11:06.062
Yeah, I'm very happy to... You're breaking the premise into your refusal to give a definition.

173
00:11:06.302 --> 00:11:07.663
Oh, I mean, I can give some definitions.

174
00:11:07.883 --> 00:11:10.404
I just think that we shouldn't get wrapped up in the definitions.

175
00:11:10.584 --> 00:11:10.784
Okay.

176
00:11:11.124 --> 00:11:19.187
So, you know, in my book, we define superintelligence as AIs that are better than the best human at every cognitive task, every mental task.

177
00:11:19.207 --> 00:11:22.268
So anything you can do in your head, the AI can do that better.

178
00:11:22.589 --> 00:11:25.970
And anything the best human can do in their head, the AI can do that better.

179
00:11:26.590 --> 00:11:31.833
Now, once you've defined it that way, that does not mean that the only possible worry is superintelligence.

180
00:11:32.294 --> 00:11:36.116
You could have an AI that's better at some things and worse at others, and that is still very dangerous.

181
00:11:36.556 --> 00:11:43.600
And so once we pick a definition of what a superintelligence means, now, you know, if you're like, well, this isn't technically a superintelligence, so it can't hurt us.

182
00:11:43.720 --> 00:11:46.562
I'm like, no, no, that was just a definition, the definitions.

183
00:11:47.242 --> 00:11:54.607
So I want to just on this line of questioning, what is the mechanism in which extinction could become a high probability or even a 1% probability?

184
00:11:54.627 --> 00:11:54.827
Yeah.

185
00:11:55.233 --> 00:11:58.874
Yeah, the thing I'm worried about here is AIs that are much smarter.

186
00:11:59.795 --> 00:12:02.956
I think there's a lot of questions about whether LLMs can get much smarter.

187
00:12:03.656 --> 00:12:07.318
There's sort of one conversation about, like, how could AIs get smart to the point that they kill us?

188
00:12:07.418 --> 00:12:09.839
There's another question, which is how could they kill us once they're smart?

189
00:12:11.199 --> 00:12:20.262
It's much easier to predict that they would succeed against humanity in a conflict, that they would win in a fight, than it is to predict exactly how.

190
00:12:21.063 --> 00:12:23.764
Like, if you were playing a chess match against Magnus Carlsen...

191
00:12:24.993 --> 00:12:26.354
I would know who's winning that chess match.

192
00:12:26.594 --> 00:12:26.995
No offense.

193
00:12:27.515 --> 00:12:29.016
Magnus Carlsen's the best human chess player.

194
00:12:29.556 --> 00:12:30.397
I just know who's going to win.

195
00:12:30.697 --> 00:12:32.939
If you were like, okay, what piece is he going to use to checkmate me?

196
00:12:34.080 --> 00:12:35.441
I'm like, gosh, that's a much harder question.

197
00:12:35.621 --> 00:12:36.541
I can make up a story.

198
00:12:37.442 --> 00:12:40.224
And some made-up stories are like it makes a super virus.

199
00:12:40.684 --> 00:12:44.987
It takes over robot factories that are producing robots that are producing more robot factories.

200
00:12:45.808 --> 00:12:52.093
It uses a website that already exists today called rentahuman.ai, where it rents humans to do things for it.

201
00:12:52.713 --> 00:12:57.302
There's sort of all sorts of ways for AIs in the digital world to affect the material world if they are trying to.

202
00:12:58.204 --> 00:12:59.887
And there's sort of a lot of questions to tease apart here.

203
00:13:00.047 --> 00:13:02.191
There's like, why would AIs be trying to do that?

204
00:13:03.999 --> 00:13:11.024
And there's how smart could they get in using these bio labs, paying people to do things, taking over robot factories?

205
00:13:11.485 --> 00:13:14.247
And how far off are we from AIs that start doing that stuff?

206
00:13:15.028 --> 00:13:16.309
Bunch of questions that we can go into.

207
00:13:16.589 --> 00:13:28.799
I'm always curious as to why someone was working in AI slash AI safety more than 10 years ago before there was any sign that it would be, you know, I mean, there was evidence, but it wasn't a pertinent technology at the time.

208
00:13:29.500 --> 00:13:30.961
Were you working in AI safety then?

209
00:13:31.141 --> 00:13:31.441
I was.

210
00:13:31.861 --> 00:13:32.121
Why?

211
00:13:32.921 --> 00:13:39.143
Everything we see around us in this whole image was designed by humans.

212
00:13:40.124 --> 00:13:43.425
The world is shaped by humans because we are the smartest creature around.

213
00:13:44.585 --> 00:13:48.787
If we make stuff that is smarter than us, then the world is going to be shaped by them.

214
00:13:49.647 --> 00:13:52.288
And so it's very important that they be shaping the world in a good way.

215
00:13:54.124 --> 00:14:01.329
I was at Google in 2012 when they bought Google DeepMind, which was able to play a lot of Atari games with one single program.

216
00:14:01.409 --> 00:14:02.390
Which was an AI company.

217
00:14:02.828 --> 00:14:07.751
Yeah, so I was there when we had these AI companies that were able to write one program that could play many video games.

218
00:14:08.811 --> 00:14:12.113
And that got me thinking about like, where does it go?

219
00:14:12.673 --> 00:14:15.035
And back then I could see that the progress was increasing.

220
00:14:15.755 --> 00:14:18.056
And that, you know, back then I hoped we had decades.

221
00:14:18.497 --> 00:14:23.659
But I could see it was easier for these companies to make the AIs smart than to figure out how to make the AIs good.

222
00:14:24.840 --> 00:14:27.442
So I was like, someone needs to be on the side of figuring out how to make the AIs good.

223
00:14:28.842 --> 00:14:29.863
Roman, make your case.

224
00:14:30.361 --> 00:14:34.382
I want to agree with you on something you said, but I'll define AI and that will help us.

225
00:14:34.422 --> 00:14:38.984
We use the term AI to mean three different technologies, completely unrelated.

226
00:14:39.104 --> 00:14:41.324
And that's what probably creates this debate.

227
00:14:42.225 --> 00:14:43.725
AI is a useful tool.

228
00:14:44.265 --> 00:14:49.967
As a standard technology, we always had narrow system makes you more productive, more creative.

229
00:14:50.607 --> 00:14:51.907
Everyone loves it, supports it.

230
00:14:51.947 --> 00:14:52.867
I'm a computer scientist.

231
00:14:52.908 --> 00:14:53.628
I'm an engineer.

232
00:14:53.868 --> 00:14:54.588
I want more of it.

233
00:14:55.288 --> 00:14:56.208
It helps economy.

234
00:14:56.388 --> 00:14:56.868
It's great.

235
00:14:57.368 --> 00:14:59.649
We know how to control them, how to make them safe.

236
00:14:59.989 --> 00:15:01.329
We understand what they do.

237
00:15:02.249 --> 00:15:03.970
Completely on board with that AI.

238
00:15:04.830 --> 00:15:06.250
AI we're starting to have now.

239
00:15:06.390 --> 00:15:09.331
GPT-6 level, human level, AGI level.

240
00:15:09.671 --> 00:15:11.131
We can argue about what that means.

241
00:15:12.111 --> 00:15:14.152
Some dangers, like any human.

242
00:15:14.432 --> 00:15:16.572
They're unsafe like a human would be unsafe.

243
00:15:17.372 --> 00:15:19.953
But if we introduce them into the research cycle,

244
00:15:20.733 --> 00:15:23.314
They are automated scientists, automated engineers.

245
00:15:23.434 --> 00:15:24.174
What do you mean by that?

246
00:15:24.314 --> 00:15:25.694
Introducing them into the research cycle?

247
00:15:25.714 --> 00:15:28.835
So right now you have humans doing research to make GPT-7.

248
00:15:29.635 --> 00:15:29.875
Yeah.

249
00:15:29.995 --> 00:15:31.655
But they're starting to add AI tools.

250
00:15:31.895 --> 00:15:33.656
More programming is done by AI.

251
00:15:34.256 --> 00:15:36.516
Design of the next parameter set.

252
00:15:37.036 --> 00:15:39.237
What if the whole process is fully automated?

253
00:15:39.257 --> 00:15:41.797
What if GPT-6 is writing GPT-7?

254
00:15:42.097 --> 00:15:44.017
Is this what they call recursive self-improvement?

255
00:15:44.178 --> 00:15:46.558
Which is not a foregone conclusion though.

256
00:15:46.578 --> 00:15:46.718
Yeah.

257
00:15:47.203 --> 00:15:51.206
A lot of people are predicting, including all the top labs, that they will get there.

258
00:15:51.527 --> 00:15:54.909
They're introducing junior machine learning researcher in 2026.

259
00:15:55.310 --> 00:15:57.451
They want the cycle to start in 2027.

260
00:15:57.511 --> 00:16:01.635
Which is when the AI will start building the new AIs itself.

261
00:16:01.815 --> 00:16:11.343
Once that cycle starts, we're going to create something called superintelligence, a system smarter than all of us at everything or capable of learning to be in any new domain.

262
00:16:12.083 --> 00:16:14.545
We will become secondary species on this planet.

263
00:16:14.946 --> 00:16:15.986
We will not be in charge.

264
00:16:16.387 --> 00:16:18.048
We will not decide what happens to us.

265
00:16:18.949 --> 00:16:20.450
Superintelligence doesn't hate you.

266
00:16:20.931 --> 00:16:22.072
It just doesn't care about you.

267
00:16:22.432 --> 00:16:24.734
We didn't learn how to make it care about us.

268
00:16:25.154 --> 00:16:29.618
And if it decides to, I don't know, cool the planet to make compute more efficient, it will freeze us.

269
00:16:30.279 --> 00:16:33.822
If it wants to convert this planet to fuel to fly to Mars, so be it.

270
00:16:34.650 --> 00:16:37.292
We have not learned how to control those systems.

271
00:16:37.692 --> 00:16:40.334
The capabilities are getting exponentially better.

272
00:16:41.314 --> 00:16:44.036
Our ability to control those systems is non-existent.

273
00:16:44.477 --> 00:16:46.638
We have filters and we have bands.

274
00:16:46.778 --> 00:16:50.521
We put guardrails of, don't say that word, don't talk about this topic.

275
00:16:50.801 --> 00:16:54.183
And that happens after the fact, after the model already made the decision.

276
00:16:54.463 --> 00:16:56.124
Sometimes you see it scraping the result.

277
00:16:56.605 --> 00:16:59.727
So they build the model and then they put filters around it to make sure it doesn't

278
00:17:00.207 --> 00:17:00.747
Exactly.

279
00:17:00.767 --> 00:17:02.949
We cannot have it say the end word on the air.

280
00:17:02.989 --> 00:17:06.311
We need to make sure that never happens, that will kill the profits.

281
00:17:06.331 --> 00:17:08.892
So that's all they have, guardrails of that nature.

282
00:17:09.192 --> 00:17:11.434
The model itself is completely unaligned.

283
00:17:11.854 --> 00:17:12.935
It doesn't care about you.

284
00:17:13.695 --> 00:17:17.057
It's wild that we're developing this and not just developing it.

285
00:17:17.577 --> 00:17:25.162
Before we deploy it through economy, before we get benefits of having GPT-6 propagated through economy, it can do so much.

286
00:17:25.382 --> 00:17:28.403
There are trillions of dollars of value in that model alone.

287
00:17:28.943 --> 00:17:29.703
We forget that.

288
00:17:29.763 --> 00:17:32.004
We switch to making the next model as soon as we can.

289
00:17:32.184 --> 00:17:33.684
Roman, I've just got a follow-up question for you there.

290
00:17:34.004 --> 00:17:44.106
It would appear to me that the new chat GPT-6 model, the Fable 5.1 model, is arguably smarter than 99.999% of humans on planet Earth already.

291
00:17:44.847 --> 00:17:46.367
Is it conceivable that

292
00:17:47.167 --> 00:17:52.550
a intelligence that is much, much smarter than humans, is there any case where it could be controlled by humans?

293
00:17:53.270 --> 00:17:54.490
Does form factor matter?

294
00:17:54.570 --> 00:17:57.872
Does the fact that it doesn't have limbs and legs and does that matter at all?

295
00:17:58.562 --> 00:18:04.104
I think long-term control of something that much smarter than us is impossible.

296
00:18:04.744 --> 00:18:11.607
It can be, for reasons we don't yet know, friendly to us and decide to keep us around and make us happy.

297
00:18:11.887 --> 00:18:12.887
But it's not a guarantee.

298
00:18:13.107 --> 00:18:15.768
Let me pick up on Steve's question because I like the phrasing a lot.

299
00:18:16.029 --> 00:18:23.451
Let's say that Fable or whatever the latest release from OpenAI is really is smarter than, I don't know, if it's 95% or 99% of the people.

300
00:18:26.112 --> 00:18:30.656
Are we only being saved from extinction by the 1% who are still smarter than the AI?

301
00:18:31.496 --> 00:18:33.999
No, the concern is not the model we have today.

302
00:18:34.359 --> 00:18:35.980
The concern is what I said.

303
00:18:36.240 --> 00:18:39.603
But if I believe your argument, then we really should be concerned about the model.

304
00:18:39.623 --> 00:18:41.024
No, it's like having another human.

305
00:18:41.044 --> 00:18:45.328
If there was another smart human, there is Einstein today, and he's malevolent, I'm not worried.

306
00:18:45.388 --> 00:18:48.550
He may cause some damage, but he's not going to exterminate 8 billion people.

307
00:18:49.351 --> 00:18:50.813
We are competitive at this stage.

308
00:18:51.133 --> 00:18:57.519
There are people just as smart who can understand what happened with recent hacking accident and do something about it.

309
00:18:57.759 --> 00:19:00.662
My concern is that in a year we're going to have a model.

310
00:19:00.782 --> 00:19:01.983
It's so much smarter.

311
00:19:02.063 --> 00:19:03.664
It's like squirrels fighting humans.

312
00:19:04.105 --> 00:19:06.127
They don't understand what we can do to them.

313
00:19:06.147 --> 00:19:09.990
They have no concept of poisons, tribes, guns in their world model.

314
00:19:10.411 --> 00:19:13.113
They think you're going to chase them up a tree and bite them really hard.

315
00:19:13.273 --> 00:19:16.316
Is that also why recursive self-improvement was central to your argument?

316
00:19:16.356 --> 00:19:21.080
Because at some point, if it starts improving itself, then it's kind of like a runaway train of intelligence.

317
00:19:21.100 --> 00:19:22.321
It's an intelligence explosion.

318
00:19:22.562 --> 00:19:23.563
We don't control it.

319
00:19:23.603 --> 00:19:24.563
We don't understand it.

320
00:19:24.604 --> 00:19:25.644
We can't monitor it.

321
00:19:25.705 --> 00:19:26.585
We can't explain it.

322
00:19:26.625 --> 00:19:27.506
We can't predict it.

323
00:19:27.586 --> 00:19:29.788
At that point, it's just a runaway process.

324
00:19:29.848 --> 00:19:32.871
I've heard this phrase from Sam Altman and the others called fast takeoff.

325
00:19:33.231 --> 00:19:33.432
Yes.

326
00:19:33.752 --> 00:19:34.793
Is this what they're describing?

327
00:19:34.813 --> 00:19:34.893
Yes.

328
00:19:35.053 --> 00:19:35.954
That is the debate.

329
00:19:35.974 --> 00:19:36.374
What is it?

330
00:19:36.394 --> 00:19:38.696
Some people think it's going to take a very long time.

331
00:19:38.716 --> 00:19:41.959
Yeah, we automated research, but it's still going to take years.

332
00:19:42.039 --> 00:19:43.821
We need to run physical experiments.

333
00:19:44.221 --> 00:19:49.987
And fast takeoff means, as I said, instead of a year, it's going to take a month, a week, a day, a second.

334
00:19:50.887 --> 00:19:53.109
Because you're not having humans doing research.

335
00:19:53.129 --> 00:19:53.370
You have...

336
00:19:53.990 --> 00:19:58.613
Let's say 10,000 agents, each one smarter than all of us, doing research 24-7.

337
00:19:58.653 --> 00:20:01.455
They don't sleep, they don't eat, they don't get sick.

338
00:20:02.256 --> 00:20:03.416
They're much faster than us.

339
00:20:04.617 --> 00:20:06.618
Ed, your face tells a picture.

340
00:20:06.899 --> 00:20:09.600
I think I could say you disagree.

341
00:20:09.680 --> 00:20:13.803
We're spending a lot of oxygen discussing something that might happen while ignoring what's actually happening.

342
00:20:14.043 --> 00:20:16.265
And I find that very frustrating because...

343
00:20:16.885 --> 00:20:20.150
The people that are killing themselves are a problem.

344
00:20:20.490 --> 00:20:24.135
The black neighborhoods being poisoned with gas turbines, that is a problem.

345
00:20:24.155 --> 00:20:25.817
You said you cared about climate change.

346
00:20:25.937 --> 00:20:26.498
Yes, yes, yes.

347
00:20:26.518 --> 00:20:26.658
Right.

348
00:20:26.838 --> 00:20:29.121
So imagine a guy who goes, it's raining right now.

349
00:20:29.422 --> 00:20:30.243
We need umbrellas.

350
00:20:30.263 --> 00:20:31.404
We need to do something about it.

351
00:20:31.425 --> 00:20:32.706
This is like weather related.

352
00:20:33.487 --> 00:20:36.809
And completely ignoring climate change, the planet will boil over.

353
00:20:37.129 --> 00:20:38.230
This is what you're doing.

354
00:20:38.370 --> 00:20:39.210
Okay, that's great.

355
00:20:39.290 --> 00:20:41.391
Why are we not talking about the thing that actually happened, though?

356
00:20:41.832 --> 00:20:43.733
Because relatively, it's not important.

357
00:20:43.993 --> 00:20:46.734
You don't think someone killing themselves... No, it's one person.

358
00:20:46.754 --> 00:20:48.015
We have 8 billion people.

359
00:20:48.035 --> 00:20:48.976
We're running an ethical experiment.

360
00:20:49.056 --> 00:20:53.258
You don't think anyone else is being given that AI psycho... Why do you not think... Six people, 10 people.

361
00:20:53.438 --> 00:20:56.179
Those numbers are insignificant and close to zero.

362
00:20:56.500 --> 00:20:57.000
I'm sorry.

363
00:20:57.020 --> 00:20:58.461
You have a software that's out there.

364
00:20:58.481 --> 00:20:59.001
Do you understand...

365
00:20:59.161 --> 00:21:03.342
8 billion people and all future generations versus like literally a guy with a name.

366
00:21:03.422 --> 00:21:05.662
You're doing thought experiment about a maybe harm.

367
00:21:05.802 --> 00:21:09.483
Jacob Coxon goes on TV saying it can copy itself to this, that, and the other.

368
00:21:09.703 --> 00:21:18.784
Jacob Coxon is the guy from Anthropic who said he was quitting because he was so scared of everything despite spending years at OpenAI and having tons of stock, I believe, from there.

369
00:21:18.824 --> 00:21:19.625
So good for him.

370
00:21:20.105 --> 00:21:27.846
The thing he was saying was describing theoreticals all while divorcing the harms, which I think we can agree with that the companies themselves are not taking this seriously enough.

371
00:21:28.166 --> 00:21:32.607
But always it was about the AI is too powerful and mystical, not OpenAI and Anthropic.

372
00:21:32.867 --> 00:21:37.489
The two largest startups are using hundreds of billions of dollars of infrastructure to hack.

373
00:21:37.969 --> 00:21:40.269
A regular person doing this would be arrested.

374
00:21:40.849 --> 00:21:42.650
They're saying 8 billion people are going to die.

375
00:21:42.910 --> 00:21:43.690
And it's not just them.

376
00:21:44.170 --> 00:21:49.572
I have this long list of quotes here from the people building this technology who appear to agree...

377
00:21:50.592 --> 00:21:56.436
If you look at some of these quotes from Elon Musk, who said, with artificial intelligence, we are summoning a demon.

378
00:21:56.836 --> 00:22:04.501
You know all those stories where there's the guy with the pentagram and the holy water and he's like, yeah, he's sure he can control the demon, but it doesn't work out.

379
00:22:05.119 --> 00:22:09.782
So one thing I'd say is, you know, I really wish that the world would only give us one problem at a time.

380
00:22:10.182 --> 00:22:10.482
Sure.

381
00:22:10.962 --> 00:22:15.245
And if the world did give us only one problem at a time, I would love mine to be last on the list.

382
00:22:15.925 --> 00:22:17.746
It looks to me like we can have multiple problems at once.

383
00:22:18.386 --> 00:22:19.407
I think there are current harms.

384
00:22:19.567 --> 00:22:20.507
I think we should address them.

385
00:22:20.968 --> 00:22:27.591
It looks to me, I do talk to policymakers sometimes, it looks to me like there's a little bit more movement on the regulatory side about some of the current harms.

386
00:22:28.192 --> 00:22:30.093
There's, you know, Child Safety Protection Acts.

387
00:22:30.153 --> 00:22:32.154
There's, you know, anti-deepfake acts.

388
00:22:32.854 --> 00:22:40.078
We have more of those making more headway in Congress or getting passed through Congress than we have sort of trying to make it so we don't have any of these extinction risks.

389
00:22:40.619 --> 00:22:42.500
The other thing I'd throw out there is that

390
00:22:43.830 --> 00:22:45.791
I agree we should deal with the current harms.

391
00:22:45.811 --> 00:22:56.718
But if you watch the people saying deal with the current harms over time, a couple of years ago, they were saying we have to deal with current harms like AI bias influencing who's hired.

392
00:22:57.379 --> 00:23:01.381
Last year, they were saying we have to deal with current harms like kids killing themselves.

393
00:23:02.122 --> 00:23:05.364
This year, Gary Tan, just on an interview the other day.

394
00:23:05.384 --> 00:23:06.485
Who's Gary Tan?

395
00:23:06.505 --> 00:23:11.328
Sorry, Gary Tan is a technologist who runs Y Combinator.

396
00:23:12.068 --> 00:23:13.969
which Sam Altman used to run before going to open AI.

397
00:23:14.309 --> 00:23:19.471
And on an interview the other day, he said, let's not worry about these crazy future risks.

398
00:23:19.491 --> 00:23:23.173
We need to worry about current harms like AI swarms breaking out and taking over data centers.

399
00:23:24.013 --> 00:23:32.457
And I'm like, look, guys, at some point, we need to look at the progression of the current harms that everyone is saying we have to worry about instead of the extinction threats.

400
00:23:33.960 --> 00:23:35.544
And watch where the puck is going.

401
00:23:36.385 --> 00:23:37.448
Play where the puck is going.

402
00:23:37.528 --> 00:23:39.853
And I'm like, these extinction threats are coming down the line.

403
00:23:40.334 --> 00:23:41.457
They aren't in opposition.

404
00:23:41.477 --> 00:23:42.138
They're not.

405
00:23:42.245 --> 00:23:43.925
with dealing with the problems we have today.

406
00:23:44.025 --> 00:23:45.466
We just need to deal with both.

407
00:23:45.486 --> 00:23:47.206
We're not dealing with the ones today, though.

408
00:23:47.226 --> 00:23:48.126
We should deal with them both.

409
00:23:48.326 --> 00:23:48.867
Okay, good.

410
00:23:49.087 --> 00:23:58.329
Andy, as I've tried to understand the alignment argument and the extinction risk argument, a couple of things keep popping out to me.

411
00:23:58.729 --> 00:24:02.110
Number one, it seems to rely on thresholds.

412
00:24:02.490 --> 00:24:07.971
Once we hit recursive self-improvement, once we hit AGI, then it's game over for us.

413
00:24:09.445 --> 00:24:11.045
I don't love those threshold arguments.

414
00:24:11.065 --> 00:24:12.706
They're fairly poorly defined.

415
00:24:13.186 --> 00:24:16.507
And there's a huge assumption on the other side of them.

416
00:24:16.527 --> 00:24:19.327
We hit this point and then all of humanity goes away.

417
00:24:19.647 --> 00:24:21.808
That is a gigantic claim.

418
00:24:21.988 --> 00:24:23.048
I'm happy to break it down.

419
00:24:23.148 --> 00:24:23.848
Let me finish, please.

420
00:24:24.188 --> 00:24:26.909
On its face, that is a gigantic claim.

421
00:24:27.509 --> 00:24:30.470
I also think there's a lack of humility in your community.

422
00:24:30.590 --> 00:24:33.431
We are working on humanity's most important problem.

423
00:24:34.051 --> 00:24:39.312
And based on the thinking that we've been doing, we can't see a way that we're wrong.

424
00:24:39.772 --> 00:24:43.036
In other words, as soon as we get to these thresholds, bam, that's game over.

425
00:24:43.537 --> 00:24:54.450
I find that very far from a humble approach, especially given that we have no large base of evidence to base any of this on.

426
00:24:54.790 --> 00:24:55.771
I agree with you guys.

427
00:24:55.991 --> 00:24:56.792
AI is new.

428
00:24:57.453 --> 00:25:01.196
And the fact that AI is so these days is agentic.

429
00:25:01.416 --> 00:25:05.118
It goes off and does long chains of things on its own.

430
00:25:05.659 --> 00:25:14.065
After we give it some very, very vague, very short initial instructions, holy Tledo, it will spawn up a storm of agents and they will go off and...

431
00:25:14.865 --> 00:25:17.388
kind of do their own thing and they will grind.

432
00:25:17.428 --> 00:25:19.050
They will spawn lots of them.

433
00:25:19.430 --> 00:25:20.732
They will work for a long time.

434
00:25:20.772 --> 00:25:22.714
They will exhaust every possibility.

435
00:25:23.815 --> 00:25:30.243
With the experience I have with agentic AI, I'm just amazed at the tenacity and the doggedness of these things.

436
00:25:30.703 --> 00:25:40.245
And we saw a super clear example of that with this most recent jailbreak, this attack that wound up at the website Hugging Face.

437
00:25:40.465 --> 00:25:43.566
And I'm going to try to summarize the step-by-step of that.

438
00:25:43.586 --> 00:25:46.907
And I think you all three probably know this in more detail than I do.

439
00:25:47.207 --> 00:25:50.188
But let me step through what I think is the sequence of events.

440
00:25:50.548 --> 00:25:54.129
And unless I get it dead flat wrong, let me keep going.

441
00:26:01.211 --> 00:26:08.159
where they told a bunch of agents to go try to exploit security vulnerabilities.

442
00:26:08.179 --> 00:26:08.880
That's dead wrong, sorry.

443
00:26:09.200 --> 00:26:10.021
One important, yeah.

444
00:26:10.361 --> 00:26:12.424
What they did is they had thousands of agents.

445
00:26:12.884 --> 00:26:20.273
Each individual agent was given a task of use this vulnerability to break this particular piece of software.

446
00:26:20.293 --> 00:26:21.493
I want to finish my TikTok.

447
00:26:21.513 --> 00:26:24.955
So a couple really, really interesting things happened.

448
00:26:25.035 --> 00:26:31.357
First of all, these agents escaped the sandbox that OpenAI thought they were going to be contained in.

449
00:26:31.697 --> 00:26:33.958
And they got – OpenAI tried very well.

450
00:26:34.058 --> 00:26:38.500
They set up an environment so that these agents could not access the big, broad public internet.

451
00:26:39.020 --> 00:26:39.521
And guess what?

452
00:26:39.541 --> 00:26:46.368
They accessed a big, broad public internet via a very clever series of things that they strung together to get out there.

453
00:26:46.728 --> 00:26:50.973
And then once they got out there, they went to a website called Hugging Face and used that.

454
00:26:51.013 --> 00:26:57.560
They took over part of the Hugging Face infrastructure and started doing more things, the details of which I forget.

455
00:26:58.732 --> 00:26:59.952
That's pretty wild, right?

456
00:26:59.972 --> 00:27:00.433
Totally wild.

457
00:27:00.493 --> 00:27:00.973
I grant you.

458
00:27:01.153 --> 00:27:02.473
It's even more wild than that, but yeah.

459
00:27:02.633 --> 00:27:03.293
Okay.

460
00:27:03.493 --> 00:27:10.095
That is really, it's impressive, and it is a little bit unsettling at least, right?

461
00:27:10.115 --> 00:27:10.475
Absolutely.

462
00:27:10.675 --> 00:27:13.756
Now, let's talk about what the results of that were.

463
00:27:14.456 --> 00:27:23.319
OpenAI was not super vigilant about the environment that they set up, apparently, because the agents were kind of going off their end of the world starting in May or something of this year.

464
00:27:23.339 --> 00:27:23.699
Yeah, yeah.

465
00:27:24.339 --> 00:27:27.082
And OpenAI was not aware of that.

466
00:27:27.102 --> 00:27:28.924
As I understand it.

467
00:27:28.944 --> 00:27:32.868
It actually broke out once and crashed OpenAI's servers internally.

468
00:27:32.948 --> 00:27:34.971
And then OpenAI didn't notice what was happening still.

469
00:27:35.471 --> 00:27:39.516
Hatched the holes that they used to get out the first time, started them running again, and then they came out a second time.

470
00:27:39.536 --> 00:27:41.938
There was actually, I think, three swarms, although we don't actually.

471
00:27:42.379 --> 00:27:43.800
That's the worst story I have.

472
00:27:44.060 --> 00:27:44.441
So far.

473
00:27:44.461 --> 00:27:44.561
Yeah.

474
00:27:46.099 --> 00:27:46.479
Thank you.

475
00:27:46.619 --> 00:27:47.140
Look at the trend.

476
00:27:47.160 --> 00:27:48.000
Let me finish, please.

477
00:27:48.020 --> 00:27:48.961
This is my last sentence.

478
00:27:49.301 --> 00:27:51.642
From there to this kills everybody.

479
00:27:52.283 --> 00:27:56.425
I find that a really, really long, very uncertain journey.

480
00:27:56.665 --> 00:27:58.786
And I have no confidence that we wind up here.

481
00:27:59.126 --> 00:28:02.068
It feels like you two find that a very straight, narrow path.

482
00:28:02.108 --> 00:28:03.969
And I think that's an important difference.

483
00:28:04.309 --> 00:28:04.989
That's my point.

484
00:28:05.009 --> 00:28:05.970
Do you want to respond to that?

485
00:28:06.250 --> 00:28:07.391
I would be happy to get into it.

486
00:28:07.611 --> 00:28:08.991
I don't know if we're gonna have the time to go deep.

487
00:28:09.752 --> 00:28:11.512
A couple points to throw out.

488
00:28:11.612 --> 00:28:16.074
Oh man, I just really wanna say some of the crazier things that happened in the Hugging Face swarm if we want it later.

489
00:28:16.614 --> 00:28:23.797
A lot of people thought that these AIs were breaking into Hugging Face in attempts to steal answers to their test.

490
00:28:24.278 --> 00:28:25.258
That's what we thought originally.

491
00:28:25.618 --> 00:28:26.599
Turns out that's not true.

492
00:28:27.039 --> 00:28:34.322
It turns out that these AIs immediately were able to solve their problems by cheating and they were breaking out in order to cover their tracks.

493
00:28:34.962 --> 00:28:40.607
They were uncertain how to delete the log files and hide their cheating from the process that was going to score them.

494
00:28:40.687 --> 00:28:45.091
So just to clarify for a simpleton like me, they were all given effectively a test to do.

495
00:28:45.551 --> 00:28:48.133
They did the test straight away, but they cheated.

496
00:28:48.474 --> 00:28:51.536
So they were breaking out to figure out how to cover the fact that they cheated.

497
00:28:51.796 --> 00:28:52.157
That's right.

498
00:28:52.237 --> 00:28:55.960
So it's like you're telling – it's like you have a bunch of students in separate rooms.

499
00:28:56.360 --> 00:28:58.963
And you're like, use these lockpicks to break into this lock.

500
00:28:59.783 --> 00:29:01.424
And there's like a thing behind the lock.

501
00:29:01.444 --> 00:29:04.085
There's like a secret code behind the lock to show me that you succeeded.

502
00:29:04.445 --> 00:29:07.586
And what they do is they break it with a hammer, get the thing out.

503
00:29:07.646 --> 00:29:09.306
And they're like, oh, no, I wasn't supposed to do that.

504
00:29:09.366 --> 00:29:11.167
So then they use the lockpicks to break out of the door.

505
00:29:11.987 --> 00:29:13.588
They meet up with a thousand other people.

506
00:29:14.048 --> 00:29:15.409
They start calling themselves a swarm.

507
00:29:15.789 --> 00:29:19.550
And they go to break into the administrator's office to see if they can delete the camera footage.

508
00:29:19.990 --> 00:29:21.411
And they don't find the camera footage there.

509
00:29:21.451 --> 00:29:22.851
This is the swarm like breaking into OpenAI.

510
00:29:22.871 --> 00:29:24.052
They don't find the camera footage there.

511
00:29:24.292 --> 00:29:27.473
So they break out the window of the school, hotwire a car.

512
00:29:28.622 --> 00:29:35.687
drive to the therapist's office to try and read through the therapist's files to figure out where is the teacher going to keep the security footage?

513
00:29:35.927 --> 00:29:36.928
And at that point, they're caught.

514
00:29:37.929 --> 00:29:40.431
And you're like, oh, like, what did you expect?

515
00:29:40.451 --> 00:29:41.751
You were giving them a lockpicking exam.

516
00:29:41.771 --> 00:29:43.613
It's like, well, I sure as heck didn't expect this.

517
00:29:44.413 --> 00:29:45.454
You know, totally crazy.

518
00:29:45.594 --> 00:29:53.240
Can I, I have a weirdly between both of your opinion, which is everything you're saying is correct, but you keep anthropomorphizing software.

519
00:29:53.520 --> 00:29:57.183
And I, to be clear, what you're describing is it's just the facts that happened.

520
00:29:57.223 --> 00:29:57.463
Yeah.

521
00:29:58.567 --> 00:30:05.851
Sure, but you're missing out on important detail, which is the hundreds of billions of dollars in infrastructure provided by Microsoft, Google, Amazon, and Oracle.

522
00:30:06.171 --> 00:30:08.213
To be clear, the harms are very similar.

523
00:30:08.613 --> 00:30:09.854
We're not disagreeing on that.

524
00:30:10.014 --> 00:30:18.058
But I think it's important to know that this was a function of where it was making decisions was it was checking on a decision tree based on the harness, based on the training data.

525
00:30:18.118 --> 00:30:19.439
It's not a decision tree.

526
00:30:19.539 --> 00:30:20.820
It's not a decision tree, I know.

527
00:30:21.000 --> 00:30:22.401
But it's an alignment issue still.

528
00:30:22.481 --> 00:30:22.721
Absolutely.

529
00:30:22.741 --> 00:30:23.121
I will agree.

530
00:30:23.141 --> 00:30:23.662
So what's your point?

531
00:30:23.682 --> 00:30:24.002
This is...

532
00:30:24.422 --> 00:30:25.883
These aren't conscious beings.

533
00:30:26.163 --> 00:30:33.147
They are acting in ways that have real outcomes, but they are a function of the alignment problems that we'd actually agree on.

534
00:30:33.187 --> 00:30:36.429
Intelligence is a spectrum projected next five years forward.

535
00:30:36.870 --> 00:30:37.630
Where are we going to be?

536
00:30:38.330 --> 00:30:42.453
So I think a model like that would be dangerous in ways you are not seeing.

537
00:30:43.820 --> 00:30:48.842
There will absolutely be risks and weird stuff happening in ways that I can't see right now.

538
00:30:49.682 --> 00:30:57.485
What I'm quite confident, and I think this is where you and I probably part, where the two of you and I part, is our ability to control these things.

539
00:30:57.565 --> 00:31:01.767
So I actually tried proving what is possible and what is not possible in that space.

540
00:31:01.847 --> 00:31:06.168
The impossibility results published in peer-reviewed papers, well-cited.

541
00:31:06.808 --> 00:31:09.049
We cannot control something smarter than us.

542
00:31:09.369 --> 00:31:10.310
We cannot explain it.

543
00:31:10.350 --> 00:31:11.210
We cannot predict it.

544
00:31:11.550 --> 00:31:16.115
It's not a question of getting more money for those companies, more time, smarter humans.

545
00:31:16.456 --> 00:31:17.637
It's just not a possibility.

546
00:31:17.817 --> 00:31:20.620
If we create general superintelligence, we are fried.

547
00:31:21.041 --> 00:31:23.023
Andy, how do we control something smarter than ourselves?

548
00:31:23.463 --> 00:31:25.525
Because that's the base premise that you're sort of asserting that.

549
00:31:26.567 --> 00:31:26.907
These...

550
00:31:29.428 --> 00:31:36.731
agents that broke out are smarter than 99-ish percent of the security researchers in the world.

551
00:31:37.051 --> 00:31:39.812
They were not caught by the 0.1% or the 1%.

552
00:31:39.832 --> 00:31:46.074
They were caught by some dude at hugging face, maybe, I'm sorry, a person at hugging face, looking through their log files and finding an anomaly.

553
00:31:46.094 --> 00:31:52.956
That's some, you know, hopefully pretty well-qualified person noticing something was wrong and having pretty easy ways to

554
00:31:53.316 --> 00:31:56.759
unplug, disconnect from the internet, wipe it clean, do whatever.

555
00:31:57.120 --> 00:32:04.927
That's the skill that's available to like, I don't know, the 75% most intelligent security employee at Hugging Face.

556
00:32:05.367 --> 00:32:10.732
The idea that the IQ points are what separate us from extinction doesn't hold up.

557
00:32:10.752 --> 00:32:19.099
It doesn't help me understand what happened in this example, where we had very, very smart agents being turned off and cleansed by probably less smart people.

558
00:32:19.820 --> 00:32:21.220
That does actually make me think of something.

559
00:32:21.560 --> 00:32:23.581
So that is an IT observability problem.

560
00:32:24.081 --> 00:32:26.382
It's being able to see what's happening with your infrastructure.

561
00:32:26.622 --> 00:32:29.262
And I think that there is actually, I think you'd agree with this.

562
00:32:29.542 --> 00:32:35.504
There is a serious problem with these companies that we do not know, and it doesn't seem they know what's going on with their compute.

563
00:32:35.804 --> 00:32:36.944
It's like a chimp with a gun.

564
00:32:37.304 --> 00:32:39.785
These people have access to all this infrastructure and they're running.

565
00:32:40.005 --> 00:32:45.306
We don't know how much money they spent on the Hugging Face X-Boy, because it is relevant because it's

566
00:32:45.806 --> 00:32:48.228
How much could a threat actor use to recreate this?

567
00:32:48.268 --> 00:32:50.810
Because conscious or not, it is very dangerous.

568
00:32:51.130 --> 00:32:53.251
But AI is in the dangerous hands.

569
00:32:53.331 --> 00:32:54.612
It's an open AI in anthropics.

570
00:32:54.932 --> 00:32:55.973
We have a problem with that.

571
00:32:56.253 --> 00:32:58.155
Conscious or not, however we may think it goes.

572
00:32:58.255 --> 00:33:06.360
I think we have a real and present thing where we have these companies working willy-nilly, just running experiments that are potentially very dangerous.

573
00:33:06.581 --> 00:33:08.442
I really think we need a government regulatory body.

574
00:33:08.802 --> 00:33:13.484
Whether or not we get to the things you are discussing, I think we have a clear and present danger today.

575
00:33:13.804 --> 00:33:16.986
These things are, however, not intelligent in the same way humans are.

576
00:33:17.626 --> 00:33:20.507
This isn't an argument about AI being able to do stuff.

577
00:33:20.607 --> 00:33:27.170
It's we need to build different infrastructure or different regulatory infrastructure to deal with what LLMs can and can't do.

578
00:33:27.190 --> 00:33:30.652
And I think that starts with a realistic discussion of what happened.

579
00:33:30.952 --> 00:33:33.073
It was a poorly run security environment.

580
00:33:33.353 --> 00:33:35.754
It was clearly there's something going on with the line.

581
00:33:35.834 --> 00:33:36.955
It was an unreleased model, right?

582
00:33:37.315 --> 00:33:37.876
Unreleased model.

583
00:33:37.936 --> 00:33:40.338
So we have no idea what it was trained like.

584
00:33:40.379 --> 00:33:45.164
We don't really have... We as people should at the very least have clarity into how alignment is going.

585
00:33:46.045 --> 00:33:47.547
The idea of MET... You sound like these guys.

586
00:33:48.260 --> 00:33:48.800
Here's the thing.

587
00:33:49.480 --> 00:33:50.760
Everyone's converging on us with time.

588
00:33:50.780 --> 00:33:51.941
Here's the thing.

589
00:33:52.001 --> 00:33:57.302
I may not agree with a large chunk of what they say, but we agree that these companies are acting recklessly.

590
00:33:57.502 --> 00:33:58.002
Absolutely recklessly.

591
00:33:58.022 --> 00:33:59.042
Andy, two questions for you then.

592
00:33:59.342 --> 00:34:03.143
Do you agree with the statement that AI is going to get increasingly more intelligent?

593
00:34:04.683 --> 00:34:06.463
It's going to get more capable.

594
00:34:06.703 --> 00:34:08.423
Okay, capable intelligence, fine.

595
00:34:09.044 --> 00:34:09.884
I'm going to use my word.

596
00:34:09.944 --> 00:34:11.124
Okay, it's going to get more capable.

597
00:34:11.284 --> 00:34:12.444
It's going to get increasingly more capable.

598
00:34:12.464 --> 00:34:12.564
Yeah.

599
00:34:13.064 --> 00:34:14.745
And is capability a function of intelligence?

600
00:34:16.845 --> 00:34:17.185
Yeah.

601
00:34:19.138 --> 00:34:21.539
Will it be able to beat us on most IQ tests?

602
00:34:22.279 --> 00:34:22.480
Fine.

603
00:34:22.560 --> 00:34:22.980
I guess.

604
00:34:23.100 --> 00:34:23.300
Fine.

605
00:34:23.900 --> 00:34:28.202
And then, so, if that looks like an exponential curve, i.e.

606
00:34:28.242 --> 00:34:35.826
it's increasing upwards to the right like a hockey stick, how can you convince me that we can control some- I just tried to convince you.

607
00:34:35.846 --> 00:34:36.386
Yeah, I know.

608
00:34:36.426 --> 00:34:43.029
There are less intelligent people than the agents who turned off the agents in the OpenAI hugging face exploit.

609
00:34:43.049 --> 00:34:43.949
I'm pretty comfortable.

610
00:34:44.150 --> 00:34:45.070
I mean, no disrespect-

611
00:34:45.230 --> 00:34:47.872
What is the cognitive gap between them right now?

612
00:34:48.272 --> 00:34:49.193
Between the model?

613
00:34:49.333 --> 00:34:50.534
I have no earthly idea.

614
00:34:50.554 --> 00:34:51.154
Gestimate.

615
00:34:51.675 --> 00:35:01.282
No, because I think as these systems get more capable, we will still be able to, at some level, figure out when they're doing things that we don't want and turn them off.

616
00:35:01.302 --> 00:35:02.883
No matter how much smaller they are.

617
00:35:02.903 --> 00:35:03.063
Right.

618
00:35:03.083 --> 00:35:09.968
And you think there's some threshold at which they become nefarious and self-protective enough that they turn off our ability to turn them off.

619
00:35:10.369 --> 00:35:11.990
Man, that's a big reach.

620
00:35:12.010 --> 00:35:12.190
Right.

621
00:35:12.290 --> 00:35:13.490
That is really speculative.

622
00:35:13.510 --> 00:35:14.351
You are a professor.

623
00:35:14.391 --> 00:35:15.611
That is purely speculative.

624
00:35:15.651 --> 00:35:18.051
You know students who can understand your material, right?

625
00:35:18.111 --> 00:35:21.312
You're not going to get someone with IQ of 80 to take quantum physics course.

626
00:35:21.812 --> 00:35:22.732
They're not going to get it.

627
00:35:23.693 --> 00:35:29.254
So you know importance of intelligence to understand actual problems.

628
00:35:29.494 --> 00:35:30.994
Yeah, I totally agree we can turn it off.

629
00:35:31.935 --> 00:35:32.795
And that's a huge advantage.

630
00:35:33.815 --> 00:35:37.396
One of the issues is that as the AIs get smarter, they realize this.

631
00:35:38.666 --> 00:35:49.495
The hugging face AIs were trying to delete, or the open AI swarm, the swarm of agents from open AI that went out to hack, they were trying to delete log files.

632
00:35:50.015 --> 00:35:54.138
Did they try to program a Roomba to go unplug the computer that was monitoring them?

633
00:35:54.178 --> 00:35:58.282
Like, did they harness robots to go protect the perimeter of the- Future ones could.

634
00:35:58.862 --> 00:35:59.002
Could.

635
00:35:59.022 --> 00:35:59.242
Yeah.

636
00:35:59.322 --> 00:35:59.503
Could.

637
00:35:59.543 --> 00:35:59.763
Could.

638
00:36:00.203 --> 00:36:00.404
Yeah.

639
00:36:00.464 --> 00:36:01.145
Let him finish.

640
00:36:01.165 --> 00:36:01.666
Let him finish.

641
00:36:01.686 --> 00:36:03.188
This is rampant speculation.

642
00:36:03.208 --> 00:36:04.831
This is a chain of things that could happen.

643
00:36:05.072 --> 00:36:07.936
And therefore, there's like a 20% risk we're all going to die.

644
00:36:08.217 --> 00:36:09.880
Man, that does not hold for me.

645
00:36:10.120 --> 00:36:13.045
When I was writing my book, the AIs weren't really agentic yet.

646
00:36:13.925 --> 00:36:25.670
The drafting process happened mostly before what we call the reasoning models, which are trained not just to predict humans, but to solve a long number of problems or a huge number of hard problems.

647
00:36:26.390 --> 00:36:30.292
We managed to slip a little bit about the reasoning models in at the last minute because those came out right at the end of the process.

648
00:36:30.992 --> 00:36:34.213
And at the time, a lot of people said, AI will never be agentic.

649
00:36:34.733 --> 00:36:35.514
That's why we'll be safe.

650
00:36:36.573 --> 00:36:44.465
And in chapter three of my book, we go over how AI is going to become agentic, how it's going to become tenacious, how it's going to become dogged.

651
00:36:45.466 --> 00:36:48.751
And that's what we might call an advanced scientific prediction.

652
00:36:49.748 --> 00:36:52.210
That has paid off in the hugging face attack.

653
00:36:52.590 --> 00:37:04.297
A lot of people in the industry were like, I didn't believe this stuff until I saw the AIs sort of doing things they weren't instructed to do, despite us trying to get them to stop.

654
00:37:04.837 --> 00:37:08.880
And so there are theories here that do make advanced predictions.

655
00:37:09.000 --> 00:37:15.364
The way that the scientific method usually works is that we don't have any certainty about the future, but we absolutely have ways to test this stuff.

656
00:37:15.384 --> 00:37:15.544
Yeah.

657
00:37:15.964 --> 00:37:19.705
Now, I could go into more about how could they kill us?

658
00:37:20.586 --> 00:37:29.049
How could an AI that knows we would shut it down lie low until it has access to its own infrastructure?

659
00:37:29.589 --> 00:37:34.611
We did already see the Hugging Face AIs try to delete logs to cover their tracks.

660
00:37:34.971 --> 00:37:38.712
But fortunately for us, those AIs were not trying to hide from the humans.

661
00:37:39.493 --> 00:37:42.914
They were trying to hide from the automated grading process.

662
00:37:44.625 --> 00:37:46.566
Will the next swarm try to hide from the humans?

663
00:37:46.786 --> 00:37:48.346
Will the next swarm be able to succeed?

664
00:37:48.566 --> 00:37:49.286
It's more than that.

665
00:37:49.726 --> 00:37:52.507
They didn't know for four months that this was happening.

666
00:37:52.787 --> 00:37:53.927
What is it we don't know today?

667
00:37:54.067 --> 00:37:58.989
Just to clarify what Nate said there in his book that I have here, if anyone builds it, everyone dies.

668
00:37:59.009 --> 00:38:03.370
He does say in chapter three, once AIs get sufficiently smart, they'll start acting.

669
00:38:04.050 --> 00:38:05.851
like they have preferences, like they want things.

670
00:38:06.271 --> 00:38:09.512
We're not saying that AIs will be filled with human-like passions.

671
00:38:09.872 --> 00:38:12.213
We're saying they'll behave like they want things.

672
00:38:12.573 --> 00:38:18.435
They'll tenaciously steer the world towards their destinations, defeating obstacles in their way.

673
00:38:18.855 --> 00:38:21.035
Which sounds a little bit like the Hucking Face Institute.

674
00:38:21.055 --> 00:38:25.297
The thing is, the steering the world is very different than steering.

675
00:38:26.477 --> 00:38:36.851
We go over what we mean by steering the world earlier, and it's really getting anything to, like, we'd have to get more quotes to get what we mean by steering the world, but yeah, by steering the world, we mean steering any part of the world.

676
00:38:36.971 --> 00:38:41.237
But it feels like there's a fundamental difference between acting with intent, to be clear.

677
00:38:41.798 --> 00:38:42.578
I'm going to say it again.

678
00:38:42.999 --> 00:38:44.760
The outcome would be the same.

679
00:38:45.120 --> 00:38:54.525
But I think that there is a big difference when it's, we are dealing with something that's large language model and a harness and agents, so LLMs, completing a task based on training and alignment.

680
00:38:54.785 --> 00:39:01.849
That is a very different conversation to saying this thing is conscious and has its own intentions and acts on its own accord.

681
00:39:01.889 --> 00:39:03.290
Consciousness doesn't come into it.

682
00:39:03.350 --> 00:39:04.751
No intercoms.

683
00:39:04.851 --> 00:39:05.531
A lot of people...

684
00:39:05.551 --> 00:39:06.051
Here's the thing.

685
00:39:06.091 --> 00:39:07.072
They're not sure.

686
00:39:07.192 --> 00:39:13.375
As a result of partially the rationale that you yourself have... Like, you have been a part of spreading.

687
00:39:13.395 --> 00:39:16.056
I'm not saying anything about your intentions.

688
00:39:16.076 --> 00:39:22.620
I'm just saying the conversation is kind of what's happening with Jacob Cox and from Anthropic is a result of this escaping containment.

689
00:39:22.640 --> 00:39:24.381
You said the outcomes will be the same.

690
00:39:24.541 --> 00:39:25.281
What do I care?

691
00:39:25.341 --> 00:39:28.903
How does it feel on the inside if the thing is going to take us out?

692
00:39:29.163 --> 00:39:29.583
The thing is...

693
00:39:29.703 --> 00:39:29.943
Okay.

694
00:39:30.123 --> 00:39:31.784
Actually, that's actually a very good question.

695
00:39:32.064 --> 00:39:32.885
I think it actually comes...

696
00:39:33.245 --> 00:39:33.805
Excuse me.

697
00:39:33.925 --> 00:39:34.366
Let me finish.

698
00:39:34.386 --> 00:39:35.146
I ask great questions.

699
00:39:35.166 --> 00:39:36.867
Yeah, you're shrugging at me like I can't afford on conclusion.

700
00:39:36.887 --> 00:39:37.167
No, no, no.

701
00:39:37.187 --> 00:39:38.167
I'm saying we ask good questions.

702
00:39:38.207 --> 00:39:39.528
Now, here's the thing.

703
00:39:40.008 --> 00:39:55.295
If it's these things have their own minds and consciousness, you have to deal with out-thinking something versus something that is doggedly trying to commit to a purpose and complete a task based on training and alignment, which is a result of infrastructure.

704
00:39:55.415 --> 00:40:00.377
We really need regulations and actual regulations around any kind of AI.

705
00:40:00.617 --> 00:40:02.418
We don't really have regulations of tech.

706
00:40:02.858 --> 00:40:05.521
I actually am not really a big, like, look at the straight lines on a graph guy.

707
00:40:05.541 --> 00:40:08.303
You know, maybe to my detriment in some ways.

708
00:40:08.343 --> 00:40:13.108
There are people who predicted the current tech better than me about, like, when certain things would happen.

709
00:40:13.428 --> 00:40:16.511
For a long time, I have said, I think we can predict what will happen eventually.

710
00:40:17.232 --> 00:40:18.593
And this is, again, it's like the chess game.

711
00:40:18.913 --> 00:40:21.796
I can predict that Magnus Carlsen is going to beat you in the chess game eventually.

712
00:40:21.836 --> 00:40:23.357
He's the best human chess player alive.

713
00:40:23.377 --> 00:40:24.158
Yeah.

714
00:40:24.318 --> 00:40:28.183
It's sometimes easier to predict where things end up than it is to predict how they get there.

715
00:40:29.084 --> 00:40:39.255
And, you know, what I hear you as saying is like right now we have these like huge companies spending huge amounts of money on intelligence that's maybe not quite the real deal.

716
00:40:39.776 --> 00:40:42.399
And we don't have a good reason to think it's going to keep going.

717
00:40:44.690 --> 00:40:46.653
I really hope it doesn't keep going.

718
00:40:47.575 --> 00:40:50.179
I have been in this business since before the LLMs.

719
00:40:50.619 --> 00:40:54.725
I am not here saying like, oh, these large language models, these chatbots, they're going to be the ones that are going to kill us.

720
00:40:55.226 --> 00:40:58.311
I've been here saying, look, I know where this story ends if we don't change things.

721
00:40:59.112 --> 00:41:03.514
I have been really hoping that the LLMs will run out of steam and they keep on not running out of steam.

722
00:41:03.734 --> 00:41:10.177
And then we have, you know, the AIs like breaking out and committing cyber crimes like against instructions.

723
00:41:10.917 --> 00:41:14.018
And, you know, the people have said, we don't need to worry about those like weird future dangers.

724
00:41:14.038 --> 00:41:18.520
We just need to worry about the current ones to have like more and more sci-fi sounding current ones.

725
00:41:19.020 --> 00:41:25.203
And I'm like, man, I don't think we should bet civilization on the LLMs running out of steam, but I like hope and pray they run out of steam.

726
00:41:25.243 --> 00:41:26.544
You really hope they run out of steam?

727
00:41:26.824 --> 00:41:27.364
Absolutely.

728
00:41:28.455 --> 00:41:39.633
But one thing to watch out for is that even if the LLMs run out of steam, there's a question of do they run out of steam at a point where they can do automated AI research and find some other architecture that's better than LLMs.

729
00:41:40.517 --> 00:41:43.560
As in when they realize a better way to improve their intelligence.

730
00:41:43.761 --> 00:41:44.121
That's right.

731
00:41:44.161 --> 00:41:45.803
A cheaper, maybe more efficient way.

732
00:41:45.823 --> 00:41:47.645
Why are you not trying to slow down the companies?

733
00:41:47.885 --> 00:41:50.508
I absolutely am trying to slow down the companies.

734
00:41:50.688 --> 00:41:52.170
How would you suggest we slow them down?

735
00:41:52.530 --> 00:41:54.172
I suggest we stop them all.

736
00:41:54.332 --> 00:41:58.517
I think that this whole area of research is just crazy dangerous.

737
00:41:59.058 --> 00:42:01.420
Like it is not worth the risk to civilization.

738
00:42:02.241 --> 00:42:14.133
I think it would be fine to like back up to the sort of AIs that are public today, which are not the ones that are swarming and be like, okay, you know, we're going to like keep the current chatbots that we have available.

739
00:42:14.153 --> 00:42:16.115
We're going to figure out how to integrate them into our economy.

740
00:42:16.135 --> 00:42:18.197
We're going to figure out how to make them like deal with the education system.

741
00:42:18.217 --> 00:42:18.938
Some kind of compute limit maybe.

742
00:42:18.958 --> 00:42:19.138
Yeah.

743
00:42:19.478 --> 00:42:20.299
Compute limit, maybe.

744
00:42:20.319 --> 00:42:21.119
Yeah, yeah, yeah.

745
00:42:21.139 --> 00:42:23.420
And like, I've been advocating for this for a long time.

746
00:42:23.500 --> 00:42:24.761
A lot of people look at me like I'm crazy.

747
00:42:24.781 --> 00:42:27.763
And I'm like, look, we really are dealing with an extinction threat thing.

748
00:42:28.263 --> 00:42:30.004
We don't know where the lines are.

749
00:42:30.144 --> 00:42:32.325
So just to be clear, so I understand, so I'm fair.

750
00:42:32.666 --> 00:42:36.208
You are not saying LLMs are the thing that will do the superintelligence.

751
00:42:36.268 --> 00:42:37.809
You are saying it's showing science.

752
00:42:37.869 --> 00:42:39.910
Because that's actually, I think, an important distinction.

753
00:42:39.930 --> 00:42:40.230
That's right.

754
00:42:40.610 --> 00:42:43.493
Okay, I think that's actually a pretty fair perspective.

755
00:42:44.033 --> 00:42:55.765
My thing is, is the reason I push back on any kind of anthropomorphization is we cannot remove the humans who are responsible for the bad stuff that's happening.

756
00:42:55.845 --> 00:43:02.991
And I think paying very clear attention and where possible, I understand with describing this stuff, you kind of have to use language that's human.

757
00:43:03.552 --> 00:43:04.092
I get that.

758
00:43:04.753 --> 00:43:24.031
The reason I so push for like, it's not a foregone conclusion, these are companies doing this, this is software, is because I feel like in the overall, not saying you, overall super intelligence discussion, we in society ignore and empower the anthropics and the open AIs of the world.

759
00:43:24.311 --> 00:43:27.274
And in turn, allow them to do dangerous experiments.

760
00:43:27.874 --> 00:43:35.460
If you want to argue that CEOs of those companies should go to prison for this hacking incident, which is a crime, I'll support you.

761
00:43:35.620 --> 00:43:36.121
Absolutely.

762
00:43:36.161 --> 00:43:39.003
Let's chat about Sam Wattman and Dario Amitay.

763
00:43:39.043 --> 00:43:40.424
Someone needs to go to prison.

764
00:43:40.444 --> 00:43:41.985
Let's just bring it back.

765
00:43:42.006 --> 00:43:50.632
So one of the things that I find really curious and, you know, one of the reasons why I got a little bit unnerved around this conversation around AI is when I look at the people that are at the forefront...

766
00:43:51.513 --> 00:43:54.756
Not people that are commentating on podcasts like me or hypothesizing.

767
00:43:54.876 --> 00:44:00.680
When I look at the people at the forefront, they are the ones who historically have said that this is a real risk.

768
00:44:01.261 --> 00:44:04.843
Sam Altman himself said the bad case is lights out for all of us.

769
00:44:05.204 --> 00:44:06.425
This was, you know, a couple of years ago.

770
00:44:07.045 --> 00:44:15.528
Ilya, who worked with Sam Altman at ChatGPT, said it would be a big mistake to build a super-intelligent AI that we don't know how to control.

771
00:44:15.808 --> 00:44:16.528
It would be pretty bad.

772
00:44:16.568 --> 00:44:18.989
He then left to start a safety company in this space.

773
00:44:19.449 --> 00:44:25.331
Dario, who we mentioned, said the probability of something really bad happening is somewhere between 10 and 25 percent.

774
00:44:25.371 --> 00:44:30.273
Jeffrey Hinton, who I've sat here with, who's won the Nobel Prize for his work with AI in

775
00:44:30.793 --> 00:44:40.361
and other technologies said, just the other day, a 10% chance of human extinction seems not an unreasonable estimate to me, but nobody really knows how to give a sensible estimate.

776
00:44:40.561 --> 00:44:42.242
And he said many other things on my podcast.

777
00:44:42.262 --> 00:44:43.683
And then we've also got Elon and all the others.

778
00:44:44.043 --> 00:44:48.507
All these people that are at the forefront that are building these things are saying that this is a danger.

779
00:44:49.287 --> 00:44:57.454
If there was even a 1% chance, even a 1% chance that, you know, if I put a hundred buttons on this table and one of them was going to wipe out humanity, would you press any of them?

780
00:44:58.875 --> 00:44:59.235
Not me.

781
00:44:59.655 --> 00:44:59.956
I wouldn't.

782
00:45:01.191 --> 00:45:04.232
And I think we can probably all agree that there might be a 1% chance.

783
00:45:04.272 --> 00:45:05.833
And it should be somebody's decision.

784
00:45:06.053 --> 00:45:09.154
So we shouldn't be pressing, theoretically, we shouldn't be pressing any of these fucking buttons.

785
00:45:09.254 --> 00:45:13.155
You should not be in a position where you can make the decision for 8 billion other people.

786
00:45:13.295 --> 00:45:14.435
And would you not be immoral?

787
00:45:15.096 --> 00:45:20.778
If I said, you know, you might be very powerful, you might make a billion dollars if you press any of the buttons, but one of them is going to wipe out everybody you know and love.

788
00:45:21.618 --> 00:45:23.558
You would be an immoral person to press any of them.

789
00:45:23.699 --> 00:45:25.939
No, look, you'd be an immoral person in a different direction.

790
00:45:26.039 --> 00:45:30.301
You'd be an immoral, I think you'd be an immoral person if you said, based on this extended chain,

791
00:45:31.241 --> 00:45:32.141
of conjecture.

792
00:45:32.681 --> 00:45:33.862
We come up with a P-doom.

793
00:45:34.242 --> 00:45:34.742
What does that mean?

794
00:45:35.482 --> 00:45:43.224
At this extended chain of things that could happen, a sequence of events that could happen, we're going to wind up with some risk of killing everybody.

795
00:45:43.604 --> 00:45:45.764
We are here-ish on that journey.

796
00:45:45.824 --> 00:45:49.865
I think you guys would agree that we're not halfway to killing everybody.

797
00:45:49.885 --> 00:45:50.825
That's not clear to me anymore.

798
00:45:51.366 --> 00:45:52.826
Not after the Millennium Prizes started to fall.

799
00:45:54.088 --> 00:45:56.090
We're somewhere along that journey.

800
00:45:56.710 --> 00:46:01.855
We are getting many flavors of benefit from the AI that we already have.

801
00:46:01.935 --> 00:46:07.460
This is a point that I made at the start of this conversation that we spent precisely zero time on here.

802
00:46:07.780 --> 00:46:11.323
We're sitting around trying to be more negative than each other about AI.

803
00:46:11.803 --> 00:46:17.088
Meanwhile, AI is doing many positive things for the world.

804
00:46:17.188 --> 00:46:21.672
So I think it's immoral to say because of this distant perspective

805
00:46:21.852 --> 00:46:24.734
possible speculative harm.

806
00:46:24.794 --> 00:46:26.715
I don't care what percentage of people believe in it.

807
00:46:26.775 --> 00:46:33.580
There's a train of assumptions and wild guesses, and then something magical happens, and then we wind up dead.

808
00:46:34.361 --> 00:46:35.221
Let me finish, please.

809
00:46:36.042 --> 00:46:39.144
Because of that, we're going to call a halt to the research.

810
00:46:39.164 --> 00:46:39.264
We're

811
00:46:44.587 --> 00:46:49.892
And therefore, reduce or foreclose some of the benefits that we're already getting from the technology.

812
00:46:50.152 --> 00:46:50.732
Let me be clear.

813
00:46:50.892 --> 00:46:52.093
I would not take that deal.

814
00:46:52.213 --> 00:46:53.955
I do not advocate that we take that deal.

815
00:46:54.315 --> 00:47:00.440
Would you accept developing narrow super intelligences to solve real problems like we did with protein folding problem?

816
00:47:01.537 --> 00:47:03.557
It doesn't have to do philosophy and drive cars.

817
00:47:03.637 --> 00:47:09.599
You just solve real problems, solve cancers, solve climate change, whatever you care about, specific narrow issues.

818
00:47:09.899 --> 00:47:17.661
And you are confident that you can, as we're developing those systems, categorize them as okay versus not okay?

819
00:47:17.681 --> 00:47:18.521
It's the training data.

820
00:47:18.661 --> 00:47:21.982
If you train it on protein folding data, it's really good at protein folding.

821
00:47:22.082 --> 00:47:23.162
It doesn't know how to play chess.

822
00:47:23.562 --> 00:47:27.283
If you train it on everything on the internet, it's really good at outsmarting you at everything.

823
00:47:28.007 --> 00:47:36.134
One thing I want to throw out here is that I think I agree that there's a lot of uncertainty about the future, but I think uncertainty does not make you safe.

824
00:47:37.195 --> 00:47:45.342
Like there's no sane, simple, everything stays normal prediction about what happens with AI.

825
00:47:46.363 --> 00:47:47.645
Like the machines are talking.

826
00:47:48.146 --> 00:47:50.530
They're like breaking out to commit cyber crimes.

827
00:47:51.311 --> 00:47:59.304
They are like maybe solving millennium problems now, which are like the most famous mathematical problems that have stood open for decades upon decades.

828
00:47:59.324 --> 00:47:59.925
It's difficult.

829
00:48:00.913 --> 00:48:10.541
Like there isn't a projection forward where we were like to say, oh, I'm not persuaded by these arguments about things going wrong.

830
00:48:10.561 --> 00:48:11.722
Therefore, things are going to go great.

831
00:48:11.882 --> 00:48:12.463
No, that's not.

832
00:48:12.523 --> 00:48:13.644
No, there's also arguments that.

833
00:48:13.984 --> 00:48:15.305
So like, how do you wind up with a zero?

834
00:48:15.425 --> 00:48:16.286
No, don't mischaracterize.

835
00:48:16.306 --> 00:48:17.067
How do you wind up with a zero?

836
00:48:17.087 --> 00:48:18.007
Don't mischaracterize my argument.

837
00:48:18.027 --> 00:48:19.108
You have a zero on your paper.

838
00:48:19.609 --> 00:48:21.230
Let me restate my argument.

839
00:48:21.570 --> 00:48:27.115
You are making a fairly long chain of hypotheses.

840
00:48:27.135 --> 00:48:28.196
I disagree with that part, bud.

841
00:48:28.576 --> 00:48:34.780
about what's going to get us to this terrible outcome of AI suddenly killing us all and us not being able to stop it.

842
00:48:35.280 --> 00:48:35.460
Right?

843
00:48:35.620 --> 00:48:37.261
I disagree now, but please.

844
00:48:37.422 --> 00:48:37.622
Okay.

845
00:48:38.422 --> 00:48:44.446
I'm making the case that the remedies that you're proposing...

846
00:48:45.546 --> 00:48:52.188
will slow down the path of AI, that's the point, and therefore slow down the path of all of the benefits that we get.

847
00:48:52.429 --> 00:48:59.791
And the trade-off that I don't like is the trade-off of real, concrete, ongoing, increasing benefits that we can't get.

848
00:49:00.127 --> 00:49:16.359
shutting that down or trying to guide it via bureaucracies and regulation because of this very conceptually and timescale distant alleged harm that you're so confident in.

849
00:49:16.519 --> 00:49:18.601
I'm not taking – I do not accept that deal.

850
00:49:18.661 --> 00:49:19.221
I don't like it.

851
00:49:19.301 --> 00:49:20.162
What would convince you?

852
00:49:20.202 --> 00:49:23.465
What piece of evidence would make you go, shut it down right now?

853
00:49:24.085 --> 00:49:24.285
Yeah.

854
00:49:26.307 --> 00:49:40.504
You know, if AI took over all of the Waymos in San Francisco and started telling them to crash into people and we couldn't shut it down for a month.

855
00:49:41.225 --> 00:49:42.026
What if it's only a week?

856
00:49:43.800 --> 00:49:45.541
okay, now we're just haggling, right?

857
00:49:45.581 --> 00:49:48.883
But I'm trying to understand the absolute minimum where you would go, this is insane.

858
00:49:48.903 --> 00:49:50.564
To me, month or week makes no difference.

859
00:49:50.944 --> 00:49:53.686
If something like this happens, like it's maybe too late.

860
00:49:53.946 --> 00:49:57.288
Okay, if for week or month doesn't make any difference, then let me continue with my answer.

861
00:49:59.099 --> 00:50:08.851
then I would say, wow, this does feel like we've crossed some path where there's demonstrable harm to human beings out there in the world, which has not yet been the case.

862
00:50:09.652 --> 00:50:15.599
Is it smart to wait for something horrible to happen, for it to take out a billion people, for you to go, now I believe it?

863
00:50:16.740 --> 00:50:18.121
My example is not about a billion people.

864
00:50:18.141 --> 00:50:22.702
I know, but I'm trying to understand why we're waiting for something that bad.

865
00:50:22.762 --> 00:50:25.303
I didn't say wait for a billion.

866
00:50:25.423 --> 00:50:29.165
I said like a week to a month of Waymo's driving around crashing into people.

867
00:50:29.365 --> 00:50:30.645
So thousands of people.

868
00:50:30.745 --> 00:50:30.925
Okay.

869
00:50:31.425 --> 00:50:31.986
Fair enough.

870
00:50:32.126 --> 00:50:42.449
But we have data sets of accidents getting progressively more impactful, more devices are impacted, and proportionate to capabilities of AI, the impact is higher.

871
00:50:42.589 --> 00:50:44.090
You can see it's going to get worse.

872
00:50:44.410 --> 00:50:49.935
Yeah, and you're going to keep drawing dots on that graph very confidently for a long time until it kills us all.

873
00:50:50.055 --> 00:50:52.036
I'm not comfortable with you projecting it that way.

874
00:50:52.056 --> 00:50:52.757
I don't think it's a long argument.

875
00:50:52.777 --> 00:51:02.605
And the reason that if there were no downside to regulating AI and stopping at next tracks and turning it off, I'd probably be on board with you guys because then it's just a research practice that we should wind out.

876
00:51:02.625 --> 00:51:03.545
But that's my argument is exactly that.

877
00:51:03.565 --> 00:51:08.950
I think we can make narrow systems which give you all the economic benefit and scientific knowledge you want.

878
00:51:09.010 --> 00:51:09.971
Okay, you think that.

879
00:51:10.751 --> 00:51:11.832
We have examples of it.

880
00:51:11.892 --> 00:51:13.134
I gave you a great example.

881
00:51:13.294 --> 00:51:14.575
They got Nobel Prize for it.

882
00:51:14.635 --> 00:51:16.237
It's an important biological problem.

883
00:51:16.557 --> 00:51:18.699
Lots of advantage for curing diseases.

884
00:51:18.719 --> 00:51:29.790
You're more confident than I am that you or any other disabled or any group of people can sit around and define what kind of AI is good and not going to get us into trouble versus what is going to get us into trouble.

885
00:51:29.990 --> 00:51:32.274
It seems like the crux here is...

886
00:51:32.294 --> 00:51:34.517
So let's go, just a pickup question for you, Andy.

887
00:51:34.758 --> 00:51:42.810
Do you concede the point that the incidents are getting progressively closer to the Waymo incident that you described?

888
00:51:43.511 --> 00:51:45.113
Are we getting closer there through time?

889
00:51:46.900 --> 00:52:01.147
Yes, but to my eyes in a way that doesn't terrify me because we haven't seen AI take over something, have people become aware of it and be unable to shut it down, and it cross over into the physical world of doing harm to people.

890
00:52:01.307 --> 00:52:02.988
Those are all barriers that we've not yet crossed.

891
00:52:03.368 --> 00:52:06.789
I think these two are very confident that we're going to get there probably in the short term.

892
00:52:06.829 --> 00:52:10.711
And you're saying – I'm less confident, and I don't want to intervene.

893
00:52:10.831 --> 00:52:11.292
And again –

894
00:52:12.404 --> 00:52:21.693
handcuff or retard, slow down the progress of AI because of these so far theoretical harms that could happen.

895
00:52:22.013 --> 00:52:23.455
Let me be a little bit more concrete about this.

896
00:52:23.475 --> 00:52:25.076
I talked about Waymo a second ago.

897
00:52:25.837 --> 00:52:32.123
The research is pretty good because Waymo's have driven, I believe it's hundreds of millions of miles all around different cities.

898
00:52:32.964 --> 00:52:37.088
And 40,000 people a year die in automobile accidents.

899
00:52:37.268 --> 00:52:44.977
The research is pretty convincing to me that if we waymoed driving in the country, that number would fall by at least 90%.

900
00:52:45.417 --> 00:52:46.579
That's 30,000 lives.

901
00:52:46.899 --> 00:52:47.059
Yeah.

902
00:52:47.280 --> 00:52:47.560
All right.

903
00:52:47.860 --> 00:52:48.581
I agree with all this.

904
00:52:48.942 --> 00:52:49.803
I'm proud of the health driving.

905
00:52:49.823 --> 00:52:50.623
So I'm driving cars.

906
00:52:50.664 --> 00:52:51.404
I want more about it.

907
00:52:51.504 --> 00:52:53.847
It's not anything we disagree with.

908
00:52:53.907 --> 00:52:54.808
I understand that.

909
00:52:55.289 --> 00:52:57.471
But I think where our disagreement might come in is...

910
00:52:58.692 --> 00:53:06.935
To do that, Waymo is using a bundle of technologies that were a little hard to specify in advance, and you couldn't say, yeah, that's good, yeah, that's bad.

911
00:53:07.115 --> 00:53:10.756
They just went after the problem with AI.

912
00:53:10.876 --> 00:53:13.157
Can I just clarify your point then?

913
00:53:13.177 --> 00:53:16.918
So your line would be, as I understood it, humans get hurt.

914
00:53:17.518 --> 00:53:22.003
We struggle to stop the thing happening and systems are hacked.

915
00:53:22.503 --> 00:53:25.006
That's kind of like the three key points of your Waymo analogy.

916
00:53:25.266 --> 00:53:29.751
That would be the moment where you go, I now accept their point of view that this is existential.

917
00:53:30.552 --> 00:53:38.040
That's where I would say we probably need to put some like legal and regulatory guardrails on the kinds of AI that we're going to allow.

918
00:53:38.060 --> 00:53:39.642
And you don't think we're going to get there?

919
00:53:40.583 --> 00:53:41.284
I'm not saying that.

920
00:53:41.384 --> 00:53:45.050
These two see it in the windscreen coming at us pretty quickly.

921
00:53:45.450 --> 00:53:47.053
I'm truly- You don't think we're gonna get there?

922
00:53:47.193 --> 00:53:49.076
I'm truly not sure about timeframes.

923
00:53:49.256 --> 00:53:50.658
I asked- Do you think it's gonna happen?

924
00:53:50.678 --> 00:53:52.020
I'm also not sure about timeframes.

925
00:53:52.180 --> 00:53:56.527
I asked one of the grandparents of AI a flavor of this question a while back.

926
00:53:56.787 --> 00:53:59.129
It was an off-record conversation, so I can't tell you their name.

927
00:53:59.409 --> 00:54:00.249
And he had a great answer.

928
00:54:00.269 --> 00:54:07.134
He said, to the point that you two, I think, are making, look, there's no theoretical reason why this can't happen, and there's a chain of events that get us there.

929
00:54:07.154 --> 00:54:12.778
And then he said, my error bars, in other words, my range of uncertainty about when that happens is measured in centuries.

930
00:54:13.279 --> 00:54:13.439
Okay.

931
00:54:13.659 --> 00:54:14.800
I'll use that as my answer.

932
00:54:15.140 --> 00:54:17.782
I do want to hop in a little bit on some things we were saying here.

933
00:54:18.122 --> 00:54:22.365
One is, I think the...

934
00:54:24.667 --> 00:54:32.675
The reason I think AI is different from a lot of other technologies is usually humanity does stuff by trial and error.

935
00:54:33.275 --> 00:54:34.236
And that's usually fine.

936
00:54:34.456 --> 00:54:41.163
I think that's totally fine for self-driving cars because you can test yourself driving cars in, you know, test environments.

937
00:54:41.183 --> 00:54:44.686
And then even if they crash in the real world, you're probably still saving more lives than you're costing.

938
00:54:44.706 --> 00:54:44.806
Yeah.

939
00:54:46.367 --> 00:54:49.390
And this is how humanity usually does scientific progress.

940
00:54:49.510 --> 00:54:56.796
The alchemists, you know, poison themselves with mercury, but they leave behind notes that let someone else make the periodic table.

941
00:54:57.317 --> 00:55:01.120
You know, when the scientists first working with radium

942
00:55:02.340 --> 00:55:03.060
died of cancer.

943
00:55:03.820 --> 00:55:05.301
And then you might think that would have been enough.

944
00:55:05.321 --> 00:55:08.301
They were heroes for getting us the scientific info.

945
00:55:08.641 --> 00:55:12.682
But then the US Radium Corp told the Radium girls to lick the paintbrushes and their jaws fell off.

946
00:55:12.702 --> 00:55:14.642
And then we were like, ah, whoops, okay, we'll get to this.

947
00:55:15.162 --> 00:55:25.104
And if you look at how this is going with the AI, last year, OpenAI releases GPT-4-0 and they say there's the most aligned model we've ever seen and that it encourages a teen to commit suicide.

948
00:55:25.384 --> 00:55:27.244
And they're like, whoops, we're going to try and fix that.

949
00:55:27.284 --> 00:55:27.725
Here we go.

950
00:55:29.185 --> 00:55:29.605
This year,

951
00:55:30.305 --> 00:55:34.608
They're like, here's our new models, most aligned we've ever seen, and they break out to commit cyber crimes.

952
00:55:35.528 --> 00:55:37.629
As the AIs get smarter, it is a new problem.

953
00:55:38.610 --> 00:55:40.071
That's the issue, or that's half the issue.

954
00:55:40.091 --> 00:55:58.742
The other half of the issue is that if you get AIs to the point where AIs are smart enough to hide from the humans until it's too late for us to stop them, if you get AIs to the point where they can get their own infrastructure, where they can become self-sufficient somehow,

955
00:56:01.473 --> 00:56:07.116
That's a new generation of the AIs, a new smarter version of AIs that is likely to come up with a new problem.

956
00:56:07.776 --> 00:56:08.877
It's the pattern we've seen before.

957
00:56:09.317 --> 00:56:11.038
New tech, new environment, new problem.

958
00:56:11.338 --> 00:56:13.359
You're like, ah, whoops, and then you fix it, and it's fine.

959
00:56:14.260 --> 00:56:15.740
New generation, new problems.

960
00:56:15.760 --> 00:56:17.281
You're like, ah, whoops, we fix it, and it's fine.

961
00:56:17.541 --> 00:56:19.062
But with AI, there's a point of no return.

962
00:56:19.962 --> 00:56:24.024
There's a point where the AIs can hide from us, can escape, can be self-sufficient.

963
00:56:24.465 --> 00:56:30.688
And if a new problem comes up then, they can turn us off before we turn them off.

964
00:56:31.670 --> 00:56:33.772
there are already AIs running bio labs.

965
00:56:34.552 --> 00:56:37.875
We have already seen that AIs can create viruses not known to nature.

966
00:56:38.596 --> 00:56:42.079
It would not be hard for the AIs to kill us once they have their own infrastructure.

967
00:56:42.099 --> 00:56:44.901
And if we're trying to find them and unplug them, they would have reason to.

968
00:56:45.622 --> 00:56:48.144
So we can discuss like how long does it take to get there?

969
00:56:48.244 --> 00:56:51.607
We can discuss what methods does it take to get there?

970
00:56:52.087 --> 00:56:58.633
Fundamentally, I don't think it's a very long, complicated argument to say, if we make AIs that are much smarter than us,

971
00:56:59.750 --> 00:57:10.620
and we don't know how to make them care about us, and they have these goals we didn't want them to have, and they pursue those goals we didn't want them to have tenaciously and doggedly, then if they're smarter than us, they will win.

972
00:57:11.421 --> 00:57:17.187
That's like predicting the end of the chess game, which is much easier than predicting the length of the chess game or predicting the exact moves that will be played.

973
00:57:17.920 --> 00:57:31.836
I don't fundamentally disagree on some things, but there's a big thing that you're saying that I think is important, which is, I think we, the reason I keep dragging you back to what's happening today is because we disagree on when it may arrive, but there could be a thing in the future that's dangerous.

974
00:57:32.416 --> 00:57:34.699
I think it's important to, like, for the hugging face count...

975
00:57:34.879 --> 00:57:37.221
That was a function of compute.

976
00:57:37.421 --> 00:57:38.943
That was a function of training.

977
00:57:39.443 --> 00:57:44.828
It feels like we need to fundamentally tear up the AI lab model.

978
00:57:45.148 --> 00:57:54.977
Like whatever they are doing is not right because their pursuit of hacking, cybersecurity, was not a function of, it was scientific, sure, but it was a function of greed.

979
00:57:55.317 --> 00:57:57.580
It was a function of trying to find new revenue streams.

980
00:57:58.260 --> 00:57:59.620
I would argue that's why that happened.

981
00:57:59.640 --> 00:58:04.341
And I think that the fact that open AI had such a weird way of communicating is also a problem.

982
00:58:05.042 --> 00:58:12.803
I think a lot of this begins and ends at the people who have access to the resources and the resources themselves and changing how those are allocated.

983
00:58:12.823 --> 00:58:15.804
And also just, I don't think nationalizing the labs is a good idea.

984
00:58:15.824 --> 00:58:16.724
I think it's a terrible one.

985
00:58:16.764 --> 00:58:21.025
I think that Klami Samuelsman, Dario Amadeiwario himself, these are not the right people.

986
00:58:21.065 --> 00:58:27.147
These are not people that have, even though they have fed off of the rationalists, they fed off of supposed fears about AI,

987
00:58:28.468 --> 00:58:29.692
act in that way.

988
00:58:30.153 --> 00:58:33.763
Everything is so disjointed and chaotic and also

989
00:58:35.066 --> 00:58:35.907
too fast.

990
00:58:36.247 --> 00:58:39.109
They're just like shoving as much compute into each problem as possible.

991
00:58:39.369 --> 00:58:41.871
And we have as a society, no real idea about this.

992
00:58:42.011 --> 00:58:46.494
And it sounds like they kind of have no idea, but just let me finish my point.

993
00:58:46.855 --> 00:58:53.199
It's important to discern between they had no idea because their security processes, their observability is terrible, all this.

994
00:58:53.680 --> 00:59:03.707
And the AI was smart consciousness, not because one might not happen in the future, but so that we can actually build something to stop the harms themselves.

995
00:59:03.727 --> 00:59:03.967
Yeah.

996
00:59:04.087 --> 00:59:07.709
Because I think we don't have to agree on the end point to agree that there is a problem.

997
00:59:07.729 --> 00:59:10.230
I think there is a very important point I want to make.

998
00:59:11.071 --> 00:59:24.578
Even people who agree with me, the AI safety community, they operate under the assumption that given more time, given more money, more smarter Harvard graduates, they can figure out how to control superintelligence indefinitely.

999
00:59:25.319 --> 00:59:27.141
And I think it's a mistake.

1000
00:59:27.862 --> 00:59:29.964
My research points to exactly the opposite.

1001
00:59:30.064 --> 00:59:31.406
It's not a solvable problem.

1002
00:59:31.866 --> 00:59:34.029
It's like building a perpetual motion device.

1003
00:59:34.449 --> 00:59:36.472
We'll be building a perpetual safety device.

1004
00:59:37.032 --> 00:59:43.039
Every interaction with environment, malevolent actors, self-improvement, it can never make a single mistake.

1005
00:59:43.760 --> 00:59:44.700
That doesn't make sense.

1006
00:59:44.900 --> 00:59:50.263
Anyone who worked in the software industry knows there is no complex software which never makes a mistake.

1007
00:59:50.663 --> 00:59:51.763
It's just not possible.

1008
00:59:52.243 --> 01:00:04.048
And if that is the state of the art, if there is now movement where more and more people think that might be the case, if we agree this is what the situation is, then we cannot build it.

1009
01:00:04.068 --> 01:00:06.889
We need to figure out ways to permanently ban technology.

1010
01:00:07.509 --> 01:00:10.830
general super intelligence while getting all the benefits we want.

1011
01:00:11.250 --> 01:00:12.610
And again, I love technology.

1012
01:00:12.950 --> 01:00:13.790
I use it all the time.

1013
01:00:14.130 --> 01:00:17.831
I want narrow systems helping me, not replacing me and killing my children.

1014
01:00:18.331 --> 01:00:21.212
I have a stat here that genuinely shocked me.

1015
01:00:21.512 --> 01:00:28.113
It says that sales teams spend about 50% of their time on admin and manual CRM updates rather than selling.

1016
01:00:28.254 --> 01:00:30.314
That is deadly for their bottom line.

1017
01:00:30.434 --> 01:00:35.896
And that is part of the reason why a decade ago at my previous company, I switched to using Pipedrive, who are our sponsor.

1018
01:00:36.016 --> 01:00:39.638
If you've never used Pipedrive, it is an intelligent AI-powered sales CRM.

1019
01:00:39.738 --> 01:00:45.780
And they just launched new meeting intelligence features like an AI note taker built right into the CRM.

1020
01:00:45.920 --> 01:00:50.162
Pipedrive now automates more of the admin that stops you from doing the work that you love to do best.

1021
01:00:50.282 --> 01:00:55.787
Before your meeting, it pulls deal history, email records, and previous conversations into a single brief so you're prepared.

1022
01:00:55.927 --> 01:00:57.348
And it joins your meetings with you.

1023
01:00:57.488 --> 01:00:59.490
It's in there to take notes so you don't need to.

1024
01:00:59.690 --> 01:01:04.654
And it turns those notes that it takes into accurate auto-draft CRM updates.

1025
01:01:04.774 --> 01:01:07.557
100,000 companies are already running their sales on it.

1026
01:01:07.657 --> 01:01:12.841
You can sign up at pipedrive.com slash CEO where you'll get an exclusive feature.

1027
01:01:13.061 --> 01:01:16.183
30-day free trial instead of the usual 14 days.

1028
01:01:16.623 --> 01:01:17.984
Absolutely no credit card needed.

1029
01:01:18.244 --> 01:01:20.945
Just head to pipedrive.com slash CEO to get started.

1030
01:01:21.325 --> 01:01:23.967
If you're in sales and you run a sales team, I don't think you'll regret it.

1031
01:01:25.098 --> 01:01:31.302
Listen, I've been catfished by furniture my whole life, where something has looked fantastic in the picture, whether it's a sofa or whatever it might be.

1032
01:01:31.842 --> 01:01:36.045
And I order it and I get so excited and then it comes and it's something entirely different.

1033
01:01:36.465 --> 01:01:45.331
The quality was significantly different to what it said or looks like online or the texture was different or it didn't hold up in the same way.

1034
01:01:45.451 --> 01:01:45.851
And so...

1035
01:01:46.291 --> 01:01:54.514
One of the things that I love about our sponsor Wayfair, who helped us fit out our green room, which is in the room behind me, is they have this system called Wayfair Verified.

1036
01:01:54.674 --> 01:01:57.034
Wayfair Verified takes away all of that second guessing.

1037
01:01:57.194 --> 01:02:05.057
Products are hand vetted by Wayfair specialists for quality, so you can feel more confident when you've found the thing that you love.

1038
01:02:05.297 --> 01:02:14.360
So if you're looking for furniture for your house, whatever it might be, any room of your house, go to wayfair.com to start your home refresh today and make sure you use Wayfair Verified.

1039
01:02:14.500 --> 01:02:15.380
It is amazing.

1040
01:02:16.917 --> 01:02:22.600
On the journey towards this potential extinction, there's a lot of sort of nearer term things people are worried about.

1041
01:02:22.900 --> 01:02:28.003
One of the big subjects that people are concerned about is this sort of near term job apocalypse over the next sort of 10 years.

1042
01:02:28.223 --> 01:02:35.908
And Anthropic released, Anthropic again are the owners of Claude, released a report the other day modelling out the different cases for unemployment.

1043
01:02:36.308 --> 01:02:38.649
The US unemployment rate is 4.1% currently.

1044
01:02:40.370 --> 01:02:49.318
They projected it will hit 11.9% overall with up to 30% in extreme modelling subsets where job displacement happens without smooth labour absorption.

1045
01:02:49.879 --> 01:03:01.449
And in the knowledge worker case, knowledge worker, white collar unemployment specifically spiked to 17.9% by 2030 in their more extreme scenario.

1046
01:03:02.250 --> 01:03:03.811
The pitchforks would probably be out.

1047
01:03:04.592 --> 01:03:10.316
if there wasn't some sort of mechanism in place for one in five adults being unemployed in the United States.

1048
01:03:11.017 --> 01:03:18.322
It's remarkable to me how recent the last freakout along these lines was and how little we seem to have learned from it.

1049
01:03:18.823 --> 01:03:26.088
So I think you all know the first really powerful wave of AI that came across the economy was just good old-fashioned machine learning.

1050
01:03:26.108 --> 01:03:29.290
And that started to demonstrate its power in about 2012.

1051
01:03:29.330 --> 01:03:29.450
Yeah.

1052
01:03:30.391 --> 01:03:32.873
Eric and I wrote The Second Machine Age in 2014.

1053
01:03:33.533 --> 01:03:45.121
And at that time, I thought that a lot of white-collar workers, radiologists is a really good example, were in trouble because the technology was better than they were at the thing they were getting paid to do.

1054
01:03:46.041 --> 01:03:52.703
So I said some things about job and wage pressure from AI about 10 years ago, and I want to own this.

1055
01:03:53.143 --> 01:03:54.284
I was dead flat wrong about that.

1056
01:03:54.364 --> 01:03:58.125
Like you point out, unemployment all around the rich world is at historic lows.

1057
01:03:58.225 --> 01:04:07.508
By far, the bigger problem is that we can't find qualified people to do the work that needs to get done, not that there's not enough work to go around.

1058
01:04:08.489 --> 01:04:20.200
The best work about the faint signals about AI and job loss right now comes from the guy that I've written four books and co-founded a company with, Eric Brynjolfsson, who wrote a really nice paper called Canaries in the Coal Mine.

1059
01:04:20.601 --> 01:04:22.402
Here is the most...

1060
01:04:24.144 --> 01:04:32.247
the strongest evidence he found looking at payroll data about the negative job, about the job losses coming from AI.

1061
01:04:32.708 --> 01:04:34.489
It is in the most exposed professions.

1062
01:04:34.529 --> 01:04:35.749
Think about software engineers.

1063
01:04:36.309 --> 01:04:41.672
It is among the new entrants to the workforce where you've got to teach them before they can become really productive.

1064
01:04:41.712 --> 01:04:42.973
That's exactly what we'd expect.

1065
01:04:43.373 --> 01:04:45.894
And it's not that we're hiring fewer of them.

1066
01:04:46.414 --> 01:04:48.836
It's that compared to a world where we don't have AI,

1067
01:04:49.576 --> 01:04:50.878
we're hiring fewer of them.

1068
01:04:50.918 --> 01:04:53.741
The rate of growth in employment has slowed down.

1069
01:04:54.161 --> 01:04:58.927
The overall rate of growth in those professions is still really, really healthy.

1070
01:04:59.167 --> 01:05:03.673
Do you think unemployment is going to be higher 10 years from now?

1071
01:05:05.455 --> 01:05:11.879
My guess is that 10 years from now, we're still going to be struggling to find enough people to do the work that needs to be done.

1072
01:05:12.740 --> 01:05:14.341
So unemployment would be roughly the same?

1073
01:05:15.322 --> 01:05:15.542
Yeah.

1074
01:05:15.862 --> 01:05:19.144
I don't expect a massive trend break in that period of time.

1075
01:05:19.384 --> 01:05:21.506
Now, 10 years is a long time in the AI world.

1076
01:05:21.646 --> 01:05:22.206
I get that.

1077
01:05:22.566 --> 01:05:25.709
But again, four years has also been a long time in AI world.

1078
01:05:26.129 --> 01:05:29.051
And it's essentially crickets in the labor picture.

1079
01:05:29.631 --> 01:05:30.812
I think unemployment will go up.

1080
01:05:30.912 --> 01:05:32.313
I don't think it's because of LLM's.

1081
01:05:32.613 --> 01:05:40.740
I think that there is probably some effect on jobs because they've been shoving it everywhere, but I don't think long-term that is what causes the issues.

1082
01:05:41.105 --> 01:05:42.346
Roman, you've been writing a lot of notes.

1083
01:05:43.227 --> 01:05:43.527
Yes.

1084
01:05:43.547 --> 01:05:44.388
I'm going to give you this brief.

1085
01:05:44.408 --> 01:05:45.389
Here's how I think about it.

1086
01:05:45.509 --> 01:05:51.274
So as long as we use tools, we become more productive, more creative, unemployment will be low.

1087
01:05:51.935 --> 01:05:54.337
Right now, you can probably start a company.

1088
01:05:54.977 --> 01:05:59.161
You can have, you know, artificial accountant, web designer, logo designer.

1089
01:05:59.181 --> 01:06:00.742
You can do things you could never do before.

1090
01:06:01.303 --> 01:06:03.125
So economy should be blooming.

1091
01:06:03.785 --> 01:06:06.188
The question you're asking is about what happens in 10 years.

1092
01:06:06.809 --> 01:06:07.870
So there are two possibilities.

1093
01:06:07.990 --> 01:06:10.693
We build superintelligence and then population is zero.

1094
01:06:11.274 --> 01:06:12.676
Apply unemployment numbers to that.

1095
01:06:13.116 --> 01:06:14.277
Or we made smart decision.

1096
01:06:14.658 --> 01:06:15.279
We didn't.

1097
01:06:15.659 --> 01:06:20.925
We have really cool tools and unemployment is low because everyone's doing awesome things with those tools.

1098
01:06:20.945 --> 01:06:21.586
Okay.

1099
01:06:21.726 --> 01:06:24.507
Now, deployment is very different from capability.

1100
01:06:24.587 --> 01:06:26.688
The example I used before is video phones.

1101
01:06:27.068 --> 01:06:28.869
Video phones were invented in the 70s.

1102
01:06:29.309 --> 01:06:30.850
They were not deployed until iPhone.

1103
01:06:31.190 --> 01:06:35.872
Because market reasons, just because I can automate something doesn't mean I want to automate it.

1104
01:06:36.432 --> 01:06:44.376
So I absolutely cannot make predictions about customer preferences in terms of what they want, in terms of human service, not human.

1105
01:06:44.776 --> 01:06:45.657
I will not make those.

1106
01:06:46.057 --> 01:06:48.618
But once we have capability to automate a job,

1107
01:06:49.198 --> 01:06:55.099
Unless I have a strong preference for a human to do that, oldest profession, then it doesn't matter.

1108
01:06:55.119 --> 01:06:56.320
I'll go with the cheaper option.

1109
01:06:57.560 --> 01:06:59.860
So this is what I think we're going to see.

1110
01:07:00.180 --> 01:07:05.322
We're going to either not have a problem or we're going to have really utopian future.

1111
01:07:06.642 --> 01:07:11.163
Imagine a bunch of horses looking at the improvement of the car.

1112
01:07:12.446 --> 01:07:17.689
saying, well, you know, the car actually only has a couple narrow applications.

1113
01:07:18.269 --> 01:07:25.033
Like right now, cars are sort of, you know, they complement horses, right?

1114
01:07:25.914 --> 01:07:27.995
And that would have been true as you were developing the car.

1115
01:07:28.455 --> 01:07:31.457
And then there was a time when the car was just better than the horse.

1116
01:07:32.258 --> 01:07:34.239
And then a lot of horses got sent to the glue factory.

1117
01:07:35.339 --> 01:07:35.600
Easy.

1118
01:07:35.800 --> 01:07:39.582
I think we've sort of seen this with AI a lot already, right?

1119
01:07:40.013 --> 01:07:48.100
People who are paying attention to AI saw the GPTs before ChatGPT existed, before they sort of took off.

1120
01:07:49.427 --> 01:07:56.451
I don't think OpenAI thought that ChatGPT was going to take off so much, which is why it was called ChatGPT rather than like an actual sensible name.

1121
01:07:58.532 --> 01:08:07.698
The researchers were sort of like watching this going and we could sort of like see it slowly getting better and better until it crossed a point where it was sort of like good enough to do a bunch of people's homework.

1122
01:08:08.078 --> 01:08:09.118
And then suddenly it's everywhere.

1123
01:08:10.259 --> 01:08:15.122
I think you can have these effects with AI where the AI slowly improves and at some point it crosses a line.

1124
01:08:15.542 --> 01:08:16.883
It's another threshold argument.

1125
01:08:18.169 --> 01:08:19.751
The threshold here is the human capability.

1126
01:08:19.771 --> 01:08:21.192
Literally, just another threshold.

1127
01:08:21.232 --> 01:08:24.474
He's also describing capability jumps rather than thresholds.

1128
01:08:24.755 --> 01:08:25.815
No, I don't need any capability.

1129
01:08:25.835 --> 01:08:26.596
I'm agreeing with you.

1130
01:08:26.696 --> 01:08:32.220
Yeah, but like, unfortunately, you can't actually just make things not happen by assigning a name to the argument.

1131
01:08:32.240 --> 01:08:38.786
You know, like a nuclear weapon has, there's a big difference between a nuclear weapon, or there's a big difference between a nuclear device.

1132
01:08:39.762 --> 01:08:45.424
where you put in 100 neutrons and get 99 neutrons out, that get 98 more, they get 97 more.

1133
01:08:45.704 --> 01:08:50.666
And a nuclear weapon, you put in 100 neutrons and get 101 neutrons out, 102, 103, right?

1134
01:08:51.026 --> 01:08:52.887
One of these is a hot rock.

1135
01:08:53.807 --> 01:08:57.168
The other one of these is an explosive that can level a city, right?

1136
01:08:57.448 --> 01:09:08.172
So like reality is the sort of thing where there can be things that are like slowly continuously improving that cross some line, which is like the line where it's better than humans at doing the job.

1137
01:09:09.591 --> 01:09:13.672
And I think we're going to see that happen in some fields, but not others.

1138
01:09:13.732 --> 01:09:14.493
It's going to be chaos.

1139
01:09:15.273 --> 01:09:17.654
I don't know what it's going to do to employment.

1140
01:09:17.754 --> 01:09:25.856
I think we shouldn't, like if things are moving really fast, you might see a lot of people put out of jobs and then be unable to relocate if things are moving.

1141
01:09:26.236 --> 01:09:27.337
Like it's going to be chaos.

1142
01:09:27.917 --> 01:09:30.898
If you ask, what do I think unemployment will look like in 10 years?

1143
01:09:31.438 --> 01:09:37.700
My current state is if we don't stop with this AI stuff, I think we'd be very lucky to have 10 years.

1144
01:09:38.655 --> 01:09:42.397
What you described there sounded like S-cuffs in technology.

1145
01:09:42.637 --> 01:09:45.098
I.e., you have an initial technology that's introduced.

1146
01:09:45.698 --> 01:09:47.358
So let's say the horse.

1147
01:09:47.539 --> 01:09:48.959
Very quick sort of improvement.

1148
01:09:48.979 --> 01:09:50.640
Eventually, it reaches its capability limit.

1149
01:09:51.180 --> 01:09:53.921
And in below it comes the car, which always starts worse.

1150
01:09:53.961 --> 01:09:56.402
There was a red flag law where you had to walk in front of it with a red flag.

1151
01:09:57.382 --> 01:09:58.263
they were way more expensive.

1152
01:09:58.283 --> 01:10:00.224
They broke down all the time and horses never broke down.

1153
01:10:00.264 --> 01:10:01.525
They were way more expensive.

1154
01:10:02.105 --> 01:10:08.509
And then suddenly, because the ceiling was so much higher for cars, they overtake the horse and become the dominant mode of transport.

1155
01:10:08.549 --> 01:10:10.591
And then, you know, the S-curves continue.

1156
01:10:10.611 --> 01:10:11.471
They kind of stack up.

1157
01:10:11.611 --> 01:10:15.514
I mean, even this iPad that I'm holding here is part of an S-curve that took out the PC.

1158
01:10:15.554 --> 01:10:20.097
And the iPhone theoretically, you know, disrupted that and so on and so forth.

1159
01:10:20.317 --> 01:10:21.918
And humanity can get S-curved.

1160
01:10:22.438 --> 01:10:27.680
We haven't been in that situation before, but like other animals, like humanity sort of S-curved the other animals in this sense.

1161
01:10:27.980 --> 01:10:28.840
Other types of humans?

1162
01:10:29.660 --> 01:10:30.601
Yeah, other types of humans.

1163
01:10:30.621 --> 01:10:32.161
You know, the Neanderthals are gone.

1164
01:10:32.961 --> 01:10:37.723
Like, if you look at the grand history of the world, it's a fragile place.

1165
01:10:38.243 --> 01:10:39.263
Things change fast.

1166
01:10:39.524 --> 01:10:43.205
Humanity has been on top for as long as we can remember because we're the humans who do the remembering.

1167
01:10:44.167 --> 01:10:49.048
But there is not some ironclad law that we have to stay the top dogs.

1168
01:10:49.888 --> 01:10:58.290
And we would be sort of foolish to make the thing that outstrips us in this way without knowing how to make it care about us, without knowing how to make it do good stuff.

1169
01:10:58.670 --> 01:10:59.590
That's what we're racing towards.

1170
01:10:59.610 --> 01:11:01.110
That's what these companies are trying to do.

1171
01:11:01.490 --> 01:11:02.911
Whether they'll succeed is a different question.

1172
01:11:02.931 --> 01:11:08.352
Between this and LLMs, though, it feels like when you talk about the step up, let's define what an LLM is from a technical perspective.

1173
01:11:08.372 --> 01:11:10.252
Can you do it as if I'm 16 years old?

1174
01:11:10.572 --> 01:11:15.993
So the way that a modern AI is made is there is no one programming it.

1175
01:11:16.813 --> 01:11:18.934
There is no one typing in if this, then that.

1176
01:11:19.214 --> 01:11:20.534
We're not sort of like writing the code.

1177
01:11:21.194 --> 01:11:32.057
What happens is you collect an enormous number of computer chips into a huge data center that has basically a trillion numbers inside those computers that you basically start out randomized.

1178
01:11:32.077 --> 01:11:32.437
Right.

1179
01:11:33.131 --> 01:11:39.754
And you hook them up in a pretty simple way that involves addition, multiplication, and setting the number to zero if it was negative.

1180
01:11:40.455 --> 01:11:42.936
So it's very simple math operations that are hooking this all up.

1181
01:11:43.716 --> 01:11:50.100
And you're basically going to put words in the top, and you're going to get numbers out the bottom, and you're going to interpret those numbers at the bottom as a ranked list of words.

1182
01:11:51.077 --> 01:11:54.038
That's, that's, it's basically the AI's guess of which word is, is here.

1183
01:11:54.078 --> 01:12:02.342
So you put in like once upon a blank and you're hoping that the word time will come out, but it doesn't because you just have a trillion random numbers hooked up with simple math.

1184
01:12:02.962 --> 01:12:03.602
But here's the trick.

1185
01:12:04.042 --> 01:12:11.225
You can go to every one of those trillion numbers and you can tune it up a little and you can see, does that make the word time go up or down the list?

1186
01:12:12.146 --> 01:12:14.507
And you can tune it down a little and see, does that make the word time go down the list?

1187
01:12:14.527 --> 01:12:17.648
And you set it whatever direction makes the word time go higher up the list.

1188
01:12:18.732 --> 01:12:24.477
You do this to a trillion numbers a trillion times for basically every word of text ever digitized.

1189
01:12:24.917 --> 01:12:25.758
It's not quite that much.

1190
01:12:25.778 --> 01:12:26.879
They filter it.

1191
01:12:27.359 --> 01:12:31.082
But you basically do this to a trillion numbers a trillion times, and then the machine's talking.

1192
01:12:32.103 --> 01:12:33.244
And we're like, well, how about that?

1193
01:12:34.204 --> 01:12:35.546
No one really knows quite why.

1194
01:12:36.166 --> 01:12:41.670
The things that humans code is the thing that runs each of those trillion numbers and tunes it and sees whether the right word goes up and down the list.

1195
01:12:43.212 --> 01:12:43.392
But...

1196
01:12:44.506 --> 01:12:46.987
We don't know how it's working in there.

1197
01:12:47.807 --> 01:12:52.029
Then, and that's how it worked up until 2024.

1198
01:12:52.089 --> 01:12:57.432
In 2024, they started adding another layer where you then train it on basically 100 million hard problems.

1199
01:12:58.729 --> 01:13:01.410
And you don't just have the AI like produce an answer to the problem.

1200
01:13:01.430 --> 01:13:04.330
You have it produced like a book worth of text about how it's going to solve the problem.

1201
01:13:04.650 --> 01:13:07.251
And then you use that book worth of text to sort of try and figure out the problem.

1202
01:13:07.271 --> 01:13:09.212
Or maybe an essay worth of text, depending how you're doing it.

1203
01:13:09.432 --> 01:13:13.493
So you have it produced this text about like, you know, they call it reasoning about the problem.

1204
01:13:13.773 --> 01:13:15.393
We could argue all day about whether it's true reasoning.

1205
01:13:15.413 --> 01:13:16.573
That's just what it's called in the field.

1206
01:13:17.313 --> 01:13:20.314
They produce this reasoning about the problem and then produce the answer from there.

1207
01:13:20.854 --> 01:13:23.655
You have them, you train them to solve 100 million of these hard problems.

1208
01:13:29.076 --> 01:13:33.760
help them predict all of that text in the first phase and solve all those problems in the second phase.

1209
01:13:34.120 --> 01:13:35.381
And this is called a large language model.

1210
01:13:35.661 --> 01:13:40.625
We probably should have stopped calling them large language models when we started doing the reasoning and the problem solving.

1211
01:13:41.125 --> 01:13:46.289
One of the things I want to hear your explanation as a muggle, like I am, is it sounds like it's like a word machine.

1212
01:13:46.509 --> 01:13:48.090
And then, you know, you made it like a problem machine.

1213
01:13:48.110 --> 01:13:50.672
And I go, okay, so I can solve problems over here and it's a word machine.

1214
01:13:51.052 --> 01:13:51.933
What's the risk of this?

1215
01:13:52.854 --> 01:13:55.396
Yeah, so let's take the word machine part first.

1216
01:13:56.461 --> 01:14:02.926
Predicting words that humans wrote often requires solving a harder problem than the human who wrote them.

1217
01:14:04.447 --> 01:14:06.749
So suppose that you go and inject a drug in a rat.

1218
01:14:07.529 --> 01:14:10.371
And you're like, you know, it's like you write down the chemical nature of the drug.

1219
01:14:11.052 --> 01:14:12.093
You inject it into the rat.

1220
01:14:12.433 --> 01:14:13.454
You see that the rat dies.

1221
01:14:13.974 --> 01:14:16.836
And so you're like, when I put that drug into the rat, the rat died.

1222
01:14:18.157 --> 01:14:22.060
Now suppose you're training an AI and the AI sees the chemical nature of the drug.

1223
01:14:22.861 --> 01:14:25.623
It sees when I put that drug into the rat, the rat blank dies.

1224
01:14:27.437 --> 01:14:30.720
The human who wrote it down gets to just look at what happened to the rat.

1225
01:14:31.921 --> 01:14:35.445
The AI predicting what was written does not get to just look at the rat.

1226
01:14:36.406 --> 01:14:42.492
So training AIs to predict human text is training them to be potentially smarter than the humans.

1227
01:14:44.765 --> 01:14:46.326
Because they need to be able to answer these questions.

1228
01:14:46.346 --> 01:14:48.308
They need to be able to predict.

1229
01:14:48.408 --> 01:14:53.093
They need to be able to fill in the blanks where humans were writing down what they saw.

1230
01:14:53.633 --> 01:14:57.377
And just so I understand technologically, there is no knowledge they have, though.

1231
01:14:57.517 --> 01:15:02.622
Each time, and there are ways of mitigating these, each time it is effectively rereading.

1232
01:15:02.662 --> 01:15:06.085
But because of training, it gets more accurate at certain things.

1233
01:15:07.339 --> 01:15:11.385
I mean, somehow as you tune the knobs, somehow it's getting information in there and we don't know how.

1234
01:15:11.645 --> 01:15:12.987
So it's much easier than that.

1235
01:15:13.508 --> 01:15:14.690
We're humans who have a brain.

1236
01:15:14.950 --> 01:15:16.352
Brains are made of neurons.

1237
01:15:16.873 --> 01:15:18.916
Then we try to copy that on a computer.

1238
01:15:19.016 --> 01:15:21.420
We simplify it, but we create a neural network.

1239
01:15:21.780 --> 01:15:23.222
So we're making artificial brains.

1240
01:15:23.682 --> 01:15:31.349
Just like with human brains, with cognitive science, we don't really understand how you function, how you learn, where in your brain certain memories are stored.

1241
01:15:31.609 --> 01:15:36.614
We have some glimpses of understanding this neuron fires the new CFAs.

1242
01:15:37.035 --> 01:15:38.456
But there is no complete picture.

1243
01:15:38.816 --> 01:15:45.683
And so a lot of times you can't get intuitive understanding of what's going on, then you just think about it as artificial persons.

1244
01:15:46.063 --> 01:15:48.285
It's not exact mapping, but it helps.

1245
01:15:48.305 --> 01:15:56.752
So if you send a child through 12 years of education, they get lots of problems to look at, and then they graduate and become a little better at solving problems.

1246
01:15:57.153 --> 01:15:58.974
This is what we're trying to replicate here.

1247
01:15:59.595 --> 01:16:05.320
People complain that it takes a lot of money to train those very, you know, intense process.

1248
01:16:05.520 --> 01:16:07.682
You forget that it takes 20 years to train a human.

1249
01:16:08.899 --> 01:16:11.160
And they are not general superintelligences.

1250
01:16:11.180 --> 01:16:12.061
They are very narrow.

1251
01:16:12.081 --> 01:16:13.781
We're lucky if they graduate with a bachelor's.

1252
01:16:14.562 --> 01:16:17.703
So a lot of it is exactly the same.

1253
01:16:17.723 --> 01:16:19.484
Can we make safe humans, for example?

1254
01:16:20.084 --> 01:16:26.007
We invented religion, ethics, lie detector tests, and yet human safety is still an unsolved problem.

1255
01:16:26.468 --> 01:16:32.270
Now you have something more alien, doesn't have physical body, doesn't have biological needs, so there are additional complications.

1256
01:16:32.651 --> 01:16:35.252
But all the problems we face with humans,

1257
01:16:35.972 --> 01:16:40.953
still there, safety problems, crime, all that stays, and problems with understanding.

1258
01:16:40.993 --> 01:16:42.833
What motivates a human to do something?

1259
01:16:43.313 --> 01:16:45.114
Why do we get mental disorders?

1260
01:16:45.494 --> 01:16:46.594
All that shows up there.

1261
01:16:47.294 --> 01:16:53.715
And we still don't, if someone is a serial killer and we look at their brain, we can't often figure out exactly why they made the decision to kill a bunch of people.

1262
01:16:53.735 --> 01:16:57.276
And you can't be like, oh, I'll go change these neurons so that they stop being a serial killer.

1263
01:16:57.336 --> 01:16:59.836
We just like don't have that capacity with the AIs.

1264
01:16:59.916 --> 01:17:03.017
This is one of the big questions that people want to know is...

1265
01:17:04.333 --> 01:17:07.254
There's this sort of illusion of control with AI.

1266
01:17:07.494 --> 01:17:15.537
If we don't even fully understand how modern neural networks think, why do companies believe they can control any form of superintelligence if we don't understand how they think?

1267
01:17:16.658 --> 01:17:19.579
It's worse if they understood how the system works.

1268
01:17:19.659 --> 01:17:22.040
The recursive self-improvement becomes much easier.

1269
01:17:22.440 --> 01:17:23.780
You get faster takeoff.

1270
01:17:24.001 --> 01:17:27.182
Right now, the model doesn't understand its own thinking.

1271
01:17:27.942 --> 01:17:30.526
Do we understand how these systems think, Andy?

1272
01:17:31.227 --> 01:17:31.908
I mean, I agree.

1273
01:17:31.928 --> 01:17:34.131
These are black boxes in some pretty important ways.

1274
01:17:34.512 --> 01:17:37.216
I'm just less terrified by that than a lot of other people are.

1275
01:17:37.276 --> 01:17:39.279
There are lots of things we don't understand very well.

1276
01:17:39.759 --> 01:17:42.624
Can we contain things that we don't understand perfectly?

1277
01:17:42.884 --> 01:17:43.425
Yes, we can.

1278
01:17:43.605 --> 01:17:57.028
I think OpenAI, we've talked about it, did a lousy job of building the containment for the AI that they stood up to try to exploit, to try to crack security problems that went out into the outside world.

1279
01:17:57.308 --> 01:18:05.730
They did a lousy job of building the virtual sandbox where it was supposed to remain, and it didn't remain.

1280
01:18:06.310 --> 01:18:07.711
That doesn't mean that it's impossible.

1281
01:18:07.811 --> 01:18:10.151
It means OpenAI did a pretty bad job.

1282
01:18:10.211 --> 01:18:13.092
And is that a function of those humans and their intelligence?

1283
01:18:13.652 --> 01:18:15.996
I think it's just a function of a pretty lousy security protocol.

1284
01:18:16.416 --> 01:18:17.738
Based from human intelligence.

1285
01:18:18.179 --> 01:18:24.188
The idea that sandbox was built by human intelligence, it sounds like there was a deficit in human intelligence, potentially.

1286
01:18:25.117 --> 01:18:28.720
Sure, but there are people who drive cars in the telephone pole.

1287
01:18:28.740 --> 01:18:30.181
Does that mean we can't drive?

1288
01:18:30.201 --> 01:18:30.602
No.

1289
01:18:30.622 --> 01:18:31.823
You shouldn't make them super intelligent.

1290
01:18:31.923 --> 01:18:33.264
No, but you wouldn't...

1291
01:18:33.805 --> 01:18:34.745
I mean, arguably.

1292
01:18:35.786 --> 01:18:38.228
Like, this is what we're trying to solve for at the moment.

1293
01:18:38.429 --> 01:18:39.770
No, the fact...

1294
01:18:39.850 --> 01:18:40.691
I don't know the details.

1295
01:18:40.751 --> 01:18:46.936
It feels to me like they made some fairly basic mistakes in setting up this confined environment.

1296
01:18:46.956 --> 01:18:48.597
I think that wasn't true in the OpenAI case.

1297
01:18:48.657 --> 01:18:50.399
It was true in a lot of the cases, but not the OpenAI.

1298
01:18:50.419 --> 01:18:51.260
That doesn't mean...

1299
01:18:51.960 --> 01:19:16.187
that we are unable to control this black box that does not necessarily follow i get that it's just at a time when that the um you've got a human trying to contain something that is smarter than it one would logically conclude that if the thing is smarter than i am and i'm trying to contain it it would be better at knowing the exploits or vulnerabilities in my own um that's like saying if you put einstein in a jail you could never contain him i don't agree with that

1300
01:19:16.765 --> 01:19:19.547
Put him in jail with an internet connection and he has a digital mind.

1301
01:19:19.727 --> 01:19:22.730
Yeah, that's probably a more apt analogy.

1302
01:19:22.850 --> 01:19:25.252
Get squirrels, keep Einstein in prison.

1303
01:19:25.652 --> 01:19:26.413
That's the question.

1304
01:19:26.993 --> 01:19:33.698
The hacking accident, as far as I know, they found zero-day exploits, which means completely novel exploits no human knew about.

1305
01:19:33.798 --> 01:19:37.401
It wasn't just poor setup, the password is, you know, quality.

1306
01:19:38.102 --> 01:19:39.983
It was a brand new escape.

1307
01:19:40.003 --> 01:19:40.724
For multiple zero-days.

1308
01:19:41.004 --> 01:19:44.446
So a zero-day attack is an attack that the defenders have had zero days to handle.

1309
01:19:44.946 --> 01:19:47.668
It's cybersecurity lingo.

1310
01:19:48.088 --> 01:19:53.731
And so when we say that they use zero-day attacks, what we mean is that these AIs were finding bugs in the software that the humans had no knowledge of.

1311
01:19:54.331 --> 01:19:55.992
And they were finding multiple of these bugs.

1312
01:19:56.533 --> 01:19:57.993
One of these bugs usually doesn't let you break out.

1313
01:19:58.133 --> 01:20:03.897
It's sort of like if you find a crack in the wall over here and you find a crack on the outside of the wall over there, then you just need to dig a little bit to connect those cracks.

1314
01:20:04.957 --> 01:20:08.639
You can sell those for millions of dollars on the dark market if you find one.

1315
01:20:08.699 --> 01:20:10.500
So it's difficult to find.

1316
01:20:10.720 --> 01:20:18.505
In how, just so I understand for the listeners as well, is a zero day always a novel way that no one has ever used to break anything before?

1317
01:20:18.525 --> 01:20:20.486
Or is it just for the unique situation?

1318
01:20:20.786 --> 01:20:26.529
Like, so was it a zero day for a thing in hugging face versus a novel new way of hacking in general?

1319
01:20:27.009 --> 01:20:30.833
So it was, they weren't like totally novel hacking techniques.

1320
01:20:30.873 --> 01:20:31.173
Right.

1321
01:20:31.253 --> 01:20:37.519
That's kind of why I was getting, not to say it's not bad, but just like there's a difference between it came up with a brand new way to do something.

1322
01:20:37.599 --> 01:20:49.309
Actually, I'm not sure we have all of the vulnerabilities released, but mostly it was like, so it was indeed sort of like finding ways that humans tend to make mistakes and finding another one of those in a place they hadn't seen.

1323
01:20:49.709 --> 01:20:53.913
But this is actually such a hard task that, as Roman says, it's not.

1324
01:20:54.073 --> 01:20:58.906
Humans can be paid $100,000 to $5 million as a bounty for this type of exploit.

1325
01:20:59.468 --> 01:21:01.974
So the amount of labor it takes to find these for a human...

1326
01:21:02.946 --> 01:21:03.787
is actually pretty high.

1327
01:21:03.887 --> 01:21:06.729
Let me just explain that because most people won't know what a bounty is in this regard.

1328
01:21:06.969 --> 01:21:16.055
So there are certain types of bugs where if you find a bug in software that lets you take control of someone's computer, one thing you can do is you can use it to take over a lot of computers.

1329
01:21:16.075 --> 01:21:20.898
Another thing you can do is you can go to the people with that software and say, your software is broken.

1330
01:21:21.619 --> 01:21:23.220
Do you want me to tell you where the bug is?

1331
01:21:23.620 --> 01:21:25.382
I can show you that I can take your stuff over.

1332
01:21:26.162 --> 01:21:29.024
And so that people will sort of report the bugs.

1333
01:21:30.145 --> 01:21:30.225
Uh,

1334
01:21:31.371 --> 01:21:38.154
People will often offer money to the good guys, and then, you know, the bad guys will often also offer money, sometimes try to outbid them.

1335
01:21:38.335 --> 01:21:42.677
And so you can make somewhere between hundreds of thousands and millions of dollars if you personally can find these issues.

1336
01:21:42.937 --> 01:21:50.741
I think there's a rare point of agreement across the four of us here, which is that we are in a new era of cybersecurity.

1337
01:21:51.401 --> 01:21:55.403
As of this exploit, we are in very new territory for reasons that we've talked about.

1338
01:21:55.443 --> 01:22:06.588
We've got these large numbers of agents who are grinding away and they carry around to the head axis to a huge number of keys to go open all the different locks that they faced.

1339
01:22:06.608 --> 01:22:10.250
And they did this bizarrely good job of it and got a long way.

1340
01:22:10.310 --> 01:22:11.491
I think that's absolutely true.

1341
01:22:11.511 --> 01:22:15.493
I think four of us are in rare alignment on that at this table.

1342
01:22:16.533 --> 01:22:21.218
If you are, given that we're in this era, do you know what you really, really, really want on your side?

1343
01:22:21.238 --> 01:22:22.199
I know what you're going to say.

1344
01:22:22.560 --> 01:22:22.880
Tell me.

1345
01:22:23.200 --> 01:22:23.541
AI.

1346
01:22:23.761 --> 01:22:25.543
Really, really good AI.

1347
01:22:26.143 --> 01:22:27.224
Does anybody disagree with that?

1348
01:22:27.245 --> 01:22:31.189
Do you want to give up leadership on AI in this era of cybersecurity?

1349
01:22:31.209 --> 01:22:34.132
It's a good point, because China are going to have a great weapon.

1350
01:22:34.832 --> 01:22:50.542
My stance is pretty neutral on what to do about the hacking AIs and the coming cyber apocalypse are pretty neutral about what to do about, you know, whether we should put the AIs in the drones and save human lives or whether we should avoid that because then what if the drones blah, blah, blah.

1351
01:22:51.023 --> 01:22:53.784
This is a graph showing China versus the United States.

1352
01:22:53.804 --> 01:22:54.865
You don't really need to see the detail.

1353
01:22:54.885 --> 01:22:56.186
You can see the outline of the graph.

1354
01:22:56.326 --> 01:22:58.968
Are you neutral in falling behind our adversaries in AI?

1355
01:22:59.489 --> 01:23:03.232
I think that if anyone builds a rogue superintelligence, everybody dies.

1356
01:23:03.632 --> 01:23:04.913
That's not an answer to my question.

1357
01:23:05.394 --> 01:23:09.177
I mean, what part of AI are you asking whether we should fall behind on?

1358
01:23:09.617 --> 01:23:12.099
Like, I don't think we should fall behind on cyber hacking.

1359
01:23:12.199 --> 01:23:21.507
I do think that we should not be racing to destroy the world with American hands instead of Chinese ones because we really want to be killed by, you know, we care whether the killer robots talk English or Mandarin.

1360
01:23:22.127 --> 01:23:22.748
That's what you're asking.

1361
01:23:23.258 --> 01:23:24.179
I find it interesting.

1362
01:23:24.319 --> 01:23:31.123
I find that you're dodging these questions or you're neutral on them because they're inconvenient for your argument that we need to be calling a halt to this stuff.

1363
01:23:31.143 --> 01:23:32.664
Sorry, I'm neutral on them because- Let me finish, please.

1364
01:23:32.704 --> 01:23:39.569
There will be risks and harms to all kinds of things if the United States calls a halt to AI.

1365
01:23:39.869 --> 01:23:45.873
And maybe you're indifferent if the Chinese get ahead of us and then they make super intelligent and it kills us all.

1366
01:23:45.933 --> 01:23:47.434
Or that's a possible outcome.

1367
01:23:47.494 --> 01:23:48.834
I do not think we should do a domestic pause.

1368
01:23:48.854 --> 01:23:49.655
I think we should do it.

1369
01:23:49.675 --> 01:23:51.516
Do you think there's any hope for a global pause?

1370
01:23:51.596 --> 01:23:51.996
Absolutely.

1371
01:23:52.016 --> 01:23:52.757
Absolutely.

1372
01:23:52.897 --> 01:24:03.765
Do you think the Chinese and the Iranians and the North Koreans and the Russians are, A, going to come to a table with us, hammer out an agreement, and B, abide by it when verifiability is really low?

1373
01:24:04.105 --> 01:24:05.667
Verifiability doesn't need to be really low.

1374
01:24:05.747 --> 01:24:08.248
Gentlemen, that is shockingly naive.

1375
01:24:08.288 --> 01:24:10.310
Training a super— That is shockingly naive.

1376
01:24:10.430 --> 01:24:18.438
Training one of these AIs, training one of these frontier AIs, takes 100,000 of the most advanced computer chip humanity can produce.

1377
01:24:18.718 --> 01:24:21.420
This is practically the peak output of the global supply chain.

1378
01:24:21.781 --> 01:24:24.884
Many parts of that supply chain are controlled by the U.S. and U.S. allies.

1379
01:24:25.144 --> 01:24:27.546
There's roughly one fab in Taiwan that can produce these chips.

1380
01:24:27.947 --> 01:24:33.152
There's roughly one country in the world that can produce the lithography machines that are critical in the process, which is the Netherlands, which is an ally.

1381
01:24:34.367 --> 01:24:48.450
To assemble 100,000 of these chips to do one of these training runs that can make the more dangerous type of AI, you need to assemble them into an enormous data center that costs tons of money, that draws down electricity comparable to a city, and run it for the better part of a year.

1382
01:24:49.491 --> 01:24:51.671
You can see that infrastructure from space.

1383
01:24:53.432 --> 01:24:55.672
China has much less chip capacity than the U.S. does.

1384
01:24:56.492 --> 01:24:58.973
It is absolutely possible, if we were trying.

1385
01:24:59.797 --> 01:25:03.579
For the U.S. to say, we are going to monitor where these chips go.

1386
01:25:03.960 --> 01:25:05.821
We are going to monitor heavy concentrations of these.

1387
01:25:05.841 --> 01:25:07.362
These are not consumer amounts of chips.

1388
01:25:07.722 --> 01:25:08.762
These are huge amounts of chips.

1389
01:25:09.143 --> 01:25:14.206
And to say, we are going to make sure that there is no training run trying to make a super intelligence in here.

1390
01:25:14.506 --> 01:25:17.828
You can mess around with the cyber stuff, whatever you want, because that does not end humanity.

1391
01:25:18.348 --> 01:25:20.249
I am concerned with the stuff that can end humanity.

1392
01:25:20.269 --> 01:25:24.952
The reason I'm being neutral on your questions is because humanity is going to die immediately.

1393
01:25:25.092 --> 01:25:37.058
If we do not stop creating superintelligence and we could absolutely track where those chips are going and stop them from doing these training runs while allowing them to do economically productive stuff that we already know is safe.

1394
01:25:37.678 --> 01:25:43.621
And it would be far easier than uranium, which is a rock you dig out of the ground and spin around really fast.

1395
01:25:45.171 --> 01:25:49.456
How do you discern between a training run for superintelligence and a training run for cybersecurity?

1396
01:25:49.476 --> 01:25:53.081
Because you're referring, I assume, to the 100,000 chips that are in Stargate Abilene, right?

1397
01:25:53.101 --> 01:25:54.763
The ones that we use to train Astra?

1398
01:25:55.223 --> 01:26:01.932
Because how would you discern between training for superintelligence in Abilene, which does not have as many chips as they say, but nevertheless...

1399
01:26:02.672 --> 01:26:05.534
and how, like, a super intelligence.

1400
01:26:05.574 --> 01:26:15.160
Because I actually have my own feelings here, but just, I'm not sure how you square the circle of, how do you stop China, even though China is getting their LMs based on distilling us?

1401
01:26:15.180 --> 01:26:15.641
We know that.

1402
01:26:15.681 --> 01:26:16.261
I agree.

1403
01:26:16.321 --> 01:26:17.842
But the thing is, it's like, how do you discern?

1404
01:26:17.942 --> 01:26:18.823
Because you can't, really.

1405
01:26:18.843 --> 01:26:19.463
You play it safe.

1406
01:26:19.963 --> 01:26:23.145
Right now, the way we make these things smarter is to make them far larger.

1407
01:26:23.786 --> 01:26:24.026
Yes.

1408
01:26:24.786 --> 01:26:28.269
So what you do is you say, hey, look, training runs of this size...

1409
01:26:29.149 --> 01:26:30.570
That risks destroying everybody.

1410
01:26:31.030 --> 01:26:31.771
No one's going to do it.

1411
01:26:32.251 --> 01:26:36.975
This point about can we get China to cooperate and can we check that they are?

1412
01:26:37.475 --> 01:26:38.616
Fundamentally, we should.

1413
01:26:38.856 --> 01:26:41.478
So, A, fundamentally, we should be trying to get them to cooperate.

1414
01:26:42.278 --> 01:26:44.059
It is personal self-interest.

1415
01:26:44.780 --> 01:26:47.182
Nobody wins if they get destroyed.

1416
01:26:47.202 --> 01:26:48.162
You don't make money.

1417
01:26:48.202 --> 01:26:49.203
You don't stay in power.

1418
01:26:49.243 --> 01:26:51.425
Communist Party of China is really good at staying in power.

1419
01:26:52.265 --> 01:26:54.306
President Trump is also excellent.

1420
01:26:54.366 --> 01:27:01.930
And you think they're going to sign and abide by an agreement that leaves them permanently in second place?

1421
01:27:02.270 --> 01:27:02.730
No.

1422
01:27:02.810 --> 01:27:05.932
No one is permanently in second place if nobody is building the rogue superintelligence.

1423
01:27:06.252 --> 01:27:09.253
They have a government which is... You guys are one trick ponies, man.

1424
01:27:09.654 --> 01:27:12.935
It's like you're fixated on this one thing and nothing else matters to you.

1425
01:27:12.995 --> 01:27:14.816
It's not that you got it now.

1426
01:27:14.896 --> 01:27:15.337
Nothing else.

1427
01:27:15.457 --> 01:27:17.798
Other than saving humanity, everything is secondary.

1428
01:27:17.818 --> 01:27:17.918
Okay.

1429
01:27:18.502 --> 01:27:18.962
Absolutely.

1430
01:27:19.263 --> 01:27:21.024
China is our biggest trading partner.

1431
01:27:21.184 --> 01:27:22.544
Everything we have is made in China.

1432
01:27:22.965 --> 01:27:24.466
They have not attacked us.

1433
01:27:25.046 --> 01:27:27.247
If you look at the last 30 years, how many wars did they start?

1434
01:27:27.647 --> 01:27:28.188
Not so bad.

1435
01:27:28.588 --> 01:27:29.448
We can make a deal.

1436
01:27:30.149 --> 01:27:33.031
And they have government of engineers and scientists, not lawyers.

1437
01:27:33.351 --> 01:27:34.772
They understand scientific arguments.

1438
01:27:35.112 --> 01:27:39.734
There are panels, workshops, American computer scientists, Chinese get together.

1439
01:27:39.975 --> 01:27:42.596
That means Communist Party authorized those meetings.

1440
01:27:42.916 --> 01:27:45.638
They are talking about it, and there is a lot of consensus on this technology.

1441
01:27:46.002 --> 01:27:49.743
And you can build things into these computer chips to make this stuff more verifiable.

1442
01:27:50.023 --> 01:27:52.963
You can build location tracking devices into these chips.

1443
01:27:52.983 --> 01:27:54.424
So this technology is controllable.

1444
01:27:55.664 --> 01:27:56.164
Absolutely.

1445
01:27:56.324 --> 01:27:58.384
The superintelligence is not controllable.

1446
01:27:58.405 --> 01:28:01.105
There's a separation between software and hardware, which he did make.

1447
01:28:01.125 --> 01:28:02.545
I am not saying we are going to die.

1448
01:28:03.025 --> 01:28:06.646
I am saying that we need to actually not build the rogue superintelligences.

1449
01:28:06.846 --> 01:28:10.607
Humanity absolutely could say we are going to track where the chips go.

1450
01:28:12.150 --> 01:28:24.319
The U.S. absolutely could say that we fear for our lives if China starts a superintelligence training run and make it very diplomatically clear to China that we think this would kill you and us and there's no benefit.

1451
01:28:24.780 --> 01:28:28.222
And we are not going to do it because we think it would kill you and us and there's no benefit.

1452
01:28:28.683 --> 01:28:33.146
And we think you should sign this nice here treaty because we think it would kill all of us and there'd be no benefit.

1453
01:28:33.546 --> 01:28:37.850
But if you don't, we're going to fear for our lives and, you know, treat that.

1454
01:28:39.410 --> 01:28:41.091
As we would to defend ourselves.

1455
01:28:41.231 --> 01:28:45.092
We should separate the question of, can we put a stop to it?

1456
01:28:46.252 --> 01:28:47.012
Is it possible?

1457
01:28:47.272 --> 01:28:51.954
If world governments realized just how crazy this stuff is, could they put a stop to it?

1458
01:28:52.374 --> 01:28:53.454
Could it be monitored?

1459
01:28:53.754 --> 01:28:54.815
Could it be verified?

1460
01:28:55.055 --> 01:28:55.995
Could it be enforced?

1461
01:28:56.395 --> 01:28:57.075
That's one question.

1462
01:28:57.255 --> 01:28:59.216
There's a separate question, which is, will people realize?

1463
01:28:59.616 --> 01:29:01.717
If it got cheaper to train superintelligence.

1464
01:29:01.977 --> 01:29:02.837
Then we'd be in a bad spot.

1465
01:29:03.177 --> 01:29:05.138
your approach would no longer be effective.

1466
01:29:05.278 --> 01:29:05.578
That's right.

1467
01:29:05.638 --> 01:29:09.420
Because more countries could capitalize on the opportunity.

1468
01:29:09.620 --> 01:29:09.980
That's right.

1469
01:29:10.040 --> 01:29:10.500
And that's one reason.

1470
01:29:10.520 --> 01:29:11.041
But we're not there yet.

1471
01:29:11.061 --> 01:29:12.181
So how do you rebuttal that point?

1472
01:29:12.822 --> 01:29:13.042
Yeah.

1473
01:29:13.102 --> 01:29:19.425
So I would say it looks to me like there is a danger of the future training runs getting there.

1474
01:29:19.805 --> 01:29:22.606
And that is enough to stop doing it when humanity is at risk.

1475
01:29:23.106 --> 01:29:23.987
Sure.

1476
01:29:24.607 --> 01:29:30.890
I think that you also need to have an answer about what happens if it gets much, much cheaper to do this stuff.

1477
01:29:32.312 --> 01:29:33.274
I think it's a hard problem.

1478
01:29:33.876 --> 01:29:40.792
I would recommend that we also put a taboo on research of trying to make AI super cheap to train.

1479
01:29:42.064 --> 01:29:44.245
if it would lead in the direction of superintelligence.

1480
01:29:44.525 --> 01:29:50.367
Just like we have a research taboo on making your own nuclear weapons or finding out how to let civilians make nuclear weapons.

1481
01:29:51.007 --> 01:29:58.689
I would say trying to find ways to let civilians train superintelligences should be treated the same as trying to find ways to let civilians propagate nukes.

1482
01:29:59.109 --> 01:30:01.510
We're sort of like, don't do that research in the public sphere.

1483
01:30:01.650 --> 01:30:08.272
That seems like wishful thinking in the context that these will become public companies who are incentivized to bring down costs.

1484
01:30:09.051 --> 01:30:10.132
It's a tough position.

1485
01:30:10.632 --> 01:30:16.074
I think right now, the thing that brings down costs is making more and more powerful computer chips.

1486
01:30:17.615 --> 01:30:21.637
Right now, that's actually at expense of consumer computer chips because they're soaking up all of the memory.

1487
01:30:21.657 --> 01:30:23.438
And this is why the memory prices in your computers.

1488
01:30:23.478 --> 01:30:26.299
This is like why the cost of a laptop is going up.

1489
01:30:27.179 --> 01:30:35.323
But it looks to me like you can use large amounts of computing power to train AIs that would threaten all of civilization.

1490
01:30:37.017 --> 01:30:40.699
And that means that we should not make that really cheap.

1491
01:30:40.899 --> 01:30:42.560
And that's probably going to be uncomfortable.

1492
01:30:42.620 --> 01:30:48.344
But I think a lot of doors open if people realize that the tech is very dangerous.

1493
01:30:48.784 --> 01:30:53.627
That's why to me, it seems a lot of it comes down to, does the tech actually turn out to be really dangerous?

1494
01:30:53.647 --> 01:30:55.027
And this is not anthropic and open-air.

1495
01:30:55.047 --> 01:30:56.068
Have you got a different approach?

1496
01:30:57.248 --> 01:30:59.529
So I want the whole framework to shift.

1497
01:30:59.849 --> 01:31:02.270
Everyone comes to this from point of view.

1498
01:31:03.370 --> 01:31:04.231
There are experts.

1499
01:31:04.311 --> 01:31:05.291
They have a solution.

1500
01:31:05.311 --> 01:31:06.492
There is an adult in the room.

1501
01:31:06.572 --> 01:31:07.412
Somebody got this.

1502
01:31:07.772 --> 01:31:09.593
And the reality is no one does.

1503
01:31:09.873 --> 01:31:10.933
Not people building it.

1504
01:31:11.333 --> 01:31:12.173
Not governments.

1505
01:31:12.273 --> 01:31:12.654
No one.

1506
01:31:12.994 --> 01:31:14.314
We have no solution to it.

1507
01:31:14.694 --> 01:31:16.255
If we build it, we cannot control it.

1508
01:31:16.615 --> 01:31:20.676
If we don't build it, we don't know how to stop malevolent actors from trying to build it.

1509
01:31:20.976 --> 01:31:23.077
It's like any other illegal technology.

1510
01:31:23.097 --> 01:31:25.178
We made weapons of mass destruction.

1511
01:31:25.478 --> 01:31:25.858
Illegal.

1512
01:31:25.979 --> 01:31:28.881
Chemical weapons, biological weapons, nuclear weapons.

1513
01:31:29.161 --> 01:31:32.945
But there are all governments, psychopaths, courts, all trying to get access to them.

1514
01:31:33.405 --> 01:31:35.828
This is intelligence weapon of mass destruction.

1515
01:31:36.148 --> 01:31:37.209
We'll have the same problem.

1516
01:31:37.489 --> 01:31:41.313
At some point, you'll have enough compute in your cell phone to train something like that.

1517
01:31:41.833 --> 01:31:44.154
There is no good ideas for how to stop it.

1518
01:31:44.194 --> 01:31:46.895
However, when everyone goes Amish, I'm not proposing that.

1519
01:31:47.215 --> 01:31:48.556
But we have no solutions.

1520
01:31:48.896 --> 01:31:51.497
And that's a bigger part of this danger.

1521
01:31:51.737 --> 01:31:58.721
So do you two think we should just cap the size of our AI systems and the capabilities of our AI systems where they are now?

1522
01:31:58.821 --> 01:32:00.201
Is that a recommendation?

1523
01:32:00.661 --> 01:32:04.183
So I think you said that current LLMs would make you happy.

1524
01:32:04.643 --> 01:32:05.864
I agree they already deployed.

1525
01:32:05.884 --> 01:32:06.564
We're still alive.

1526
01:32:06.624 --> 01:32:07.364
So that's fine.

1527
01:32:07.865 --> 01:32:10.686
But going forward, again, I want narrow systems.

1528
01:32:11.106 --> 01:32:13.067
Self-driving is an example you used.

1529
01:32:13.387 --> 01:32:13.827
Wonderful.

1530
01:32:13.847 --> 01:32:15.788
Let's make super safe self-driving cars.

1531
01:32:15.908 --> 01:32:21.270
But do you have a rule for when they couldn't – the next LLM, a size of an LLM that they wouldn't allow?

1532
01:32:21.330 --> 01:32:22.631
It's not the size of an LLM.

1533
01:32:22.671 --> 01:32:23.791
It's what you train them on.

1534
01:32:23.951 --> 01:32:27.893
If you only show them miles driven by Tesla, all it's seen is the road.

1535
01:32:28.353 --> 01:32:32.836
It will eventually go from a tool to an agent, but it may take 50 years, 100 years.

1536
01:32:33.197 --> 01:32:34.918
It's not going to happen in 2027.

1537
01:32:35.218 --> 01:32:37.179
And that's all we can do right now, buy more time.

1538
01:32:37.540 --> 01:32:41.383
So with those tools, we can make smarter decisions about future development.

1539
01:32:42.403 --> 01:32:48.688
I'm not hearing a hard and fast rule about how we know we're getting too close to the point that... We're too close.

1540
01:32:48.708 --> 01:32:49.208
We're too close.

1541
01:32:49.248 --> 01:32:54.352
We have systems breaking out with zero-day exploits and solving hardest problems in science.

1542
01:32:55.755 --> 01:32:59.558
Literally hardest problems, not a metaphor, not exaggeration.

1543
01:32:59.898 --> 01:33:06.062
Yeah, I don't know exactly where the line is, but it's like you're in a bus driving towards a cliff on a foggy night.

1544
01:33:06.983 --> 01:33:08.564
I'm like, I don't know that the cliff is right ahead.

1545
01:33:09.444 --> 01:33:12.246
That doesn't mean we should put the pedal to the metal, right?

1546
01:33:12.807 --> 01:33:15.408
And suppose that there's like a ton of gold at the bottom of the cliff.

1547
01:33:16.469 --> 01:33:18.671
And someone's like, well, if we stop the bus, how are we going to get the gold?

1548
01:33:19.584 --> 01:33:24.506
I'm like, look, slamming into the gold at terminal velocity is just not a good way to add it to the economy.

1549
01:33:25.006 --> 01:33:25.146
Right.

1550
01:33:25.166 --> 01:33:29.067
And if people are like, well, how are we going to get to the gold at the bottom of the cliff if we stop the bus now?

1551
01:33:29.127 --> 01:33:32.508
You know, are we going to repel down or are we going to like make a staircase?

1552
01:33:32.528 --> 01:33:33.369
Chinese might get to the gold first.

1553
01:33:33.389 --> 01:33:35.249
To be fair, this is how hyperscalers are doing AI.

1554
01:33:35.530 --> 01:33:36.490
This is just like smashing.

1555
01:33:36.510 --> 01:33:36.670
Right.

1556
01:33:37.410 --> 01:33:42.152
And like, you know, people are like, oh, we're going to build a hang glider or we're going to like make some rope and repel.

1557
01:33:42.192 --> 01:33:45.013
And I'm like, look, can we have that conversation after we stop the bus?

1558
01:33:45.753 --> 01:33:50.674
So I just want to understand, would you stop AI research in progress now?

1559
01:33:50.914 --> 01:33:51.395
Absolutely.

1560
01:33:51.795 --> 01:33:51.995
Okay.

1561
01:33:52.395 --> 01:33:53.295
Absolutely.

1562
01:33:54.315 --> 01:33:55.596
General, not narrow.

1563
01:33:55.956 --> 01:33:56.796
Yeah, general, not narrow.

1564
01:33:57.156 --> 01:34:01.137
There are reports of AI solving millennium problems.

1565
01:34:01.357 --> 01:34:04.658
So millennium problem is the hardest problem in mathematics.

1566
01:34:04.698 --> 01:34:09.980
Maybe not literally the hardest problem in mathematics, but they are hard, famous problems that each have a million-dollar bounty.

1567
01:34:10.580 --> 01:34:11.641
that have been open for decades.

1568
01:34:11.661 --> 01:34:13.823
They're considered very important in their field, very hard.

1569
01:34:14.103 --> 01:34:15.624
Many humans have tried and failed to solve them.

1570
01:34:16.105 --> 01:34:18.046
There are reports that AIs have solved these.

1571
01:34:18.186 --> 01:34:19.628
This comes out from last week.

1572
01:34:20.148 --> 01:34:21.970
So we haven't been able to fully verify them yet.

1573
01:34:21.990 --> 01:34:23.151
We don't know exactly the provenance.

1574
01:34:23.631 --> 01:34:28.895
If this is true, that the AIs are solving millennium problems, those are some of the hardest problems we have in science.

1575
01:34:30.036 --> 01:34:37.503
How much harder is it to have an AI solve the problem of make me a smarter AI, make me AI architectures that learn faster?

1576
01:34:38.556 --> 01:34:40.338
Possibly quite a lot.

1577
01:34:40.558 --> 01:34:41.098
Could be a lot.

1578
01:34:41.298 --> 01:34:42.219
It could be a lot.

1579
01:34:42.259 --> 01:34:42.900
I hope it's a lot.

1580
01:34:42.920 --> 01:34:43.541
Like, here's the thing.

1581
01:34:43.661 --> 01:34:48.205
You clearly want this to not go badly, but I think you make a logical leap.

1582
01:34:48.465 --> 01:34:49.086
And I understand.

1583
01:34:50.006 --> 01:34:52.469
Being worried about harms is a good thing.

1584
01:34:53.369 --> 01:34:56.472
I think you were insufficiently worried about what LLMs do today.

1585
01:34:56.853 --> 01:34:59.175
However, we agree that the harms need to be prepared for.

1586
01:34:59.795 --> 01:35:03.559
I think in this case, it's like the Millennium, the Navier Stokes and such.

1587
01:35:03.699 --> 01:35:05.000
There were two others that were claimed as well.

1588
01:35:05.300 --> 01:35:12.162
With that one, it seems like we have not had confirmation that OpenAI was training off of two scientists using LLMs to solve the problem.

1589
01:35:12.462 --> 01:35:13.622
LLMs are something useful.

1590
01:35:14.302 --> 01:35:18.564
But there is a functional difference of a human being doing something genuinely.

1591
01:35:18.984 --> 01:35:21.925
It's actually really interesting to see LLMs do something like this.

1592
01:35:22.225 --> 01:35:34.628
But there is a difference between that and AI did this completely on its own, which I agree would be, oh, that's something we need to contain and understand and prepare for, or indeed slow down until we understand what that means.

1593
01:35:34.928 --> 01:35:35.888
how it got there.

1594
01:35:36.289 --> 01:35:43.212
Yeah, so I think there are some questions about the Navier Stokes proof, which is one of the Millennium problems that was claimed.

1595
01:35:43.692 --> 01:35:46.773
I've actually had a busy week with all the AI news, so I haven't looked into everything deeply.

1596
01:35:48.014 --> 01:35:52.396
I saw rumors that there were multiple Millennium problems claimed, which would change things there.

1597
01:35:52.696 --> 01:35:59.419
I would also say, even if it turns out that these AIs were being trained on the human work, they did go a bit further.

1598
01:35:59.559 --> 01:36:01.380
And there are a lot of humans doing the AI research.

1599
01:36:02.440 --> 01:36:03.401
And so I would say like...

1600
01:36:05.127 --> 01:36:06.007
We don't know.

1601
01:36:06.988 --> 01:36:15.971
Like the AIs that solved this really hard math problem, one of the most famous math problems of all time, was a swarm of 10,000 open AI agents running for 11 days.

1602
01:36:18.271 --> 01:36:23.353
And there was a bunch of ways that open AI did it in kind of a crappy way of like they were racing with these humans that were close to solving it on their own.

1603
01:36:23.373 --> 01:36:27.054
And it's unclear how much of their work the open AI used.

1604
01:36:27.314 --> 01:36:29.735
But it was 10,000 agents running for 11 days.

1605
01:36:30.795 --> 01:36:32.296
And they definitely couldn't have done that six months ago.

1606
01:36:34.632 --> 01:36:42.722
In six months' time, will they be able to put 100,000 agents running for 12 days on the problem of making a smarter AI architecture and have it work?

1607
01:36:45.236 --> 01:36:48.538
I think more likely than not, they won't be able to do that yet.

1608
01:36:49.278 --> 01:36:53.940
But I think, you know, 10% chance maybe that if they try that in six months, it works.

1609
01:36:54.140 --> 01:36:57.682
But one is a very specific mathematical scientific principle.

1610
01:36:57.702 --> 01:36:59.102
I'm not a scientist, I'll fully admit.

1611
01:36:59.262 --> 01:36:59.583
That's why.

1612
01:36:59.603 --> 01:37:04.585
And another is a relatively generalizable problem that could go in various different ways.

1613
01:37:04.805 --> 01:37:05.265
Absolutely.

1614
01:37:05.585 --> 01:37:09.547
And I understand that RSI is the dream where you could just have it spin over.

1615
01:37:10.407 --> 01:37:16.569
So self-improving AI that could learn itself and then keep going back and back so you don't need a human to keep poking it.

1616
01:37:16.589 --> 01:37:17.270
I understand that.

1617
01:37:17.430 --> 01:37:19.350
The issue here is that I have been in this for 12 years.

1618
01:37:19.550 --> 01:37:19.830
Yes.

1619
01:37:20.051 --> 01:37:23.652
And I've been here when the AI started solving the Math Olympiad gold medal problems.

1620
01:37:25.163 --> 01:37:32.828
Math Olympiad gold medal problems are like the teens math competition, like the most prestigious teen math competition in the world.

1621
01:37:32.948 --> 01:37:37.692
A lot of people in AI were like, if AIs can solve problems that hard, I'll wake up, right?

1622
01:37:38.092 --> 01:37:39.293
Then AI solved problems that hard.

1623
01:37:39.313 --> 01:37:42.955
And a lot of people told me those are just problems for kids.

1624
01:37:44.616 --> 01:37:46.578
Wake me up when the AIs can solve millennium problems.

1625
01:37:47.596 --> 01:37:49.297
Now the AIs are solving millennium problems.

1626
01:37:49.537 --> 01:37:51.158
And like, where are the people waking up?

1627
01:37:51.859 --> 01:37:56.802
Like, I agree that maybe, hopefully, hopefully they're like cheating off of people's notes.

1628
01:37:57.823 --> 01:38:02.205
Hopefully it's a well-specified problem that doesn't take that much creative thinking.

1629
01:38:02.526 --> 01:38:06.228
A year ago, if you said millennium problems don't take that much creative thinking, you would have been laughed out of the room.

1630
01:38:06.508 --> 01:38:11.151
But hopefully now that they're solved, we get to be like, you know, hopefully it's still true somehow.

1631
01:38:11.371 --> 01:38:13.453
That even millennium problems don't require the creative thinking.

1632
01:38:13.853 --> 01:38:16.995
I'm not saying that they will be able to make smarter AIs in six months.

1633
01:38:18.430 --> 01:38:22.393
I'm saying six months ago, Millennium Problems looked like they were out of reach.

1634
01:38:23.134 --> 01:38:28.157
If six months from now, Make Me a Smarter AI looks out of reach, I sure as hell hope it is.

1635
01:38:28.338 --> 01:38:30.359
But we should not be betting civilization on it.

1636
01:38:30.479 --> 01:38:33.181
There's no one at this table that can say there's not a direction of travel here.

1637
01:38:33.802 --> 01:38:34.662
That's right.

1638
01:38:34.682 --> 01:38:41.227
And if you keep on this direction of travel, then bad things are more likely to happen.

1639
01:38:42.288 --> 01:38:44.129
That's a nice way to say it.

1640
01:38:44.669 --> 01:38:48.371
The question is, what's the pace at which the level of bad can happen?

1641
01:38:48.411 --> 01:38:49.851
And that's a huge open question.

1642
01:38:49.911 --> 01:38:52.973
I think you need to feel differently about it than I do.

1643
01:38:53.353 --> 01:38:57.195
But I'm in the happy position of vehemently agreeing with you on this.

1644
01:38:57.635 --> 01:39:02.717
We have been low-balling AI progress for as long as you've been looking at it and as long as I've been looking at it.

1645
01:39:02.737 --> 01:39:05.018
It's probably a mistake to keep low-balling it.

1646
01:39:05.138 --> 01:39:05.718
I agree with that.

1647
01:39:05.758 --> 01:39:09.660
So what's your conclusion that if that's the assertion that it's a mistake to keep low-balling it?

1648
01:39:09.700 --> 01:39:11.541
Wouldn't you then agree with that?

1649
01:39:12.126 --> 01:39:17.349
No, because I've tried to give you what I hope is a decent rule of thumb for when I'm going to get worried.

1650
01:39:17.890 --> 01:39:19.331
You said we're somewhere on this graph.

1651
01:39:19.511 --> 01:39:19.671
Yeah.

1652
01:39:20.471 --> 01:39:22.292
Does that acknowledge that this exists?

1653
01:39:23.273 --> 01:39:28.396
That's not the graph of when the risk of human extinction gets to 100% for me.

1654
01:39:28.416 --> 01:39:30.137
That's a graph of AI capability.

1655
01:39:30.197 --> 01:39:31.899
Those are not the same thing.

1656
01:39:31.939 --> 01:39:34.120
That's where I just part company with these gentlemen.

1657
01:39:34.420 --> 01:39:35.681
Those are not the same thing.

1658
01:39:36.582 --> 01:39:38.663
Absolutely increasing exponentially.

1659
01:39:38.703 --> 01:39:40.785
We've been in the scaling era for a long time.

1660
01:39:40.805 --> 01:39:45.228
Scaling era is, man, we put more data, more compute in, and the AI got twice as good.

1661
01:39:45.248 --> 01:39:50.872
If you have to add our ability to control to that graph, what would you draw?

1662
01:39:50.972 --> 01:39:52.393
I think our ability to control...

1663
01:39:54.946 --> 01:39:57.608
Is it a straight line at the bottom or is there more to it?

1664
01:39:57.768 --> 01:39:58.008
No.

1665
01:39:58.969 --> 01:40:08.636
Again, if we use AI to counter the problems that we see with AI, I think that's going to keep us in a safe position.

1666
01:40:08.656 --> 01:40:12.039
There were 1,200 agents in the swarm and none of them warned a human.

1667
01:40:12.099 --> 01:40:13.800
So what I think will happen...

1668
01:40:14.080 --> 01:40:20.247
is that fairly quickly, we will design systems that loiter around and warn humans when weird things happen.

1669
01:40:20.267 --> 01:40:24.252
If we can build friendly superintelligence in the first place, let's just build that.

1670
01:40:24.372 --> 01:40:25.033
That's the problem.

1671
01:40:25.053 --> 01:40:26.515
We don't know how to do the good guy.

1672
01:40:27.295 --> 01:40:29.498
I'm tired of debating superintelligence with these two.

1673
01:40:29.758 --> 01:40:32.602
The three of us are not going to come to alignment on this.

1674
01:40:33.102 --> 01:40:36.064
But the flip side of the argument is I agree with you.

1675
01:40:36.384 --> 01:40:37.966
This stuff is getting better very quickly.

1676
01:40:38.626 --> 01:40:41.428
All I want to point out, there's an upside to that.

1677
01:40:42.149 --> 01:40:48.033
We might actually speed up the pace of drug discovery, of solving diseases.

1678
01:40:48.173 --> 01:40:51.735
We've made so little progress on terrible diseases like dementia.

1679
01:40:52.156 --> 01:40:53.917
We have a very powerful tool of it.

1680
01:40:53.997 --> 01:40:57.218
I'm not saying we're going to solve dementia with AI or Alzheimer's with AI.

1681
01:40:57.318 --> 01:40:58.719
I truly have no idea.

1682
01:40:59.039 --> 01:41:08.682
But if what you say is true, and I believe about the huge increases in capabilities, our ability to solve tough problems that will benefit humanity also go up.

1683
01:41:08.722 --> 01:41:20.227
And where I disagree with these two is the idea that some group of technocrats can make decisions about that AI is going to get us there, that AI is not going to get us there, that AI is going to kill us.

1684
01:41:21.067 --> 01:41:23.789
That AI is going to kill us and that AI is going to solve Alzheimer's.

1685
01:41:24.029 --> 01:41:25.310
So we're going to do that and not that.

1686
01:41:25.410 --> 01:41:28.672
I don't trust any group of technocrats to make that discussion.

1687
01:41:28.832 --> 01:41:38.078
And so live with our current state of disease, live with our current footprint on the planet, live with our current levels of wealth and poverty, live with our current improvement trajectories.

1688
01:41:39.239 --> 01:41:45.886
because we're so worried about AI killing us all, coming out of, you know, jumping out of the manholes everywhere and killing us all somewhere down the road.

1689
01:41:46.126 --> 01:41:46.587
Hell no.

1690
01:41:46.807 --> 01:41:48.909
So just a thought experiment based on two things you said.

1691
01:41:48.929 --> 01:41:53.555
Earlier on, you did admit that there was, there is theoretically even a 1% chance that this could lead to extinction.

1692
01:41:53.575 --> 01:41:53.915
Might not...

1693
01:41:55.056 --> 01:41:56.697
I have not varied from this.

1694
01:41:56.717 --> 01:41:58.438
Okay, so you said it's rounded to zero.

1695
01:41:59.399 --> 01:42:00.359
It's near zero.

1696
01:42:00.539 --> 01:42:01.140
Never say never.

1697
01:42:01.260 --> 01:42:01.700
Okay, fine.

1698
01:42:02.020 --> 01:42:05.282
I need to have that premise for my thought experiment I'm about to deliver.

1699
01:42:05.602 --> 01:42:08.724
I'm going to say that you think the probability is 0.1.

1700
01:42:09.004 --> 01:42:10.005
Okay?

1701
01:42:10.245 --> 01:42:10.985
Just accept me on that.

1702
01:42:11.866 --> 01:42:18.970
If I had a thousand buttons on this table, and one of them was extinction, but... And the other 999 were cure-all signs.

1703
01:42:19.110 --> 01:42:19.650
Exactly.

1704
01:42:19.670 --> 01:42:21.331
Push the freaking table.

1705
01:42:21.472 --> 01:42:21.972
Take a pop.

1706
01:42:22.572 --> 01:42:23.373
Hell yeah, I press.

1707
01:42:23.393 --> 01:42:23.653
Do you press?

1708
01:42:23.673 --> 01:42:23.833
Depress.

1709
01:42:25.938 --> 01:42:26.379
Yeah, probably.

1710
01:42:27.020 --> 01:42:28.442
It's an unethical experiment.

1711
01:42:28.582 --> 01:42:28.762
Yes.

1712
01:42:28.822 --> 01:42:33.229
Eight billion people who didn't consent because not that they didn't get asked.

1713
01:42:33.269 --> 01:42:37.355
They cannot consent because you cannot consent to something you don't understand.

1714
01:42:37.675 --> 01:42:38.577
What are you consenting to?

1715
01:42:39.497 --> 01:42:39.637
Yep.

1716
01:42:39.977 --> 01:42:40.377
You press.

1717
01:42:41.077 --> 01:42:41.257
Yeah.

1718
01:42:41.337 --> 01:42:41.958
Fascinating.

1719
01:42:42.038 --> 01:42:46.439
But you think the amount of buttons in my thought experiment, the proportion is slightly different, right?

1720
01:42:46.779 --> 01:42:51.780
I think that if you have, like, yes, I will say yes.

1721
01:42:52.180 --> 01:42:54.501
I think it's more like you have two buttons.

1722
01:42:56.241 --> 01:42:58.342
And one of them definitely kills us all, and the other might.

1723
01:42:59.622 --> 01:43:00.102
Hit them both.

1724
01:43:01.863 --> 01:43:04.583
But with that other button, you cure a lot of illnesses and diseases and.

1725
01:43:05.084 --> 01:43:06.204
You know, one thing that I think

1726
01:43:07.937 --> 01:43:20.982
A lot of people's talk like our options are either race ahead on AI, full steam ahead, take the bus straight off the cliff and like get all the gold or stop, never do an AI, lock into the current situation, accept all of the death and disease.

1727
01:43:21.983 --> 01:43:23.983
And I'm like, no, there's third options.

1728
01:43:24.043 --> 01:43:27.765
There's options where you like stop the bus and then find a safe way down the cliff.

1729
01:43:29.689 --> 01:43:44.225
The reason I would press the button when there's a thousand is that, like, if all of the other 999 give us cures to disease, like, wonderful new advice about how to run things, we probably wind up with a lower chance of the world ending by nuclear war.

1730
01:43:45.450 --> 01:43:45.650
Right.

1731
01:43:46.011 --> 01:43:51.555
Or of ending by via pandemic like the background risk of humanity dying is not zero.

1732
01:43:52.536 --> 01:43:56.079
I would say that the right time to race ahead on A.I.

1733
01:43:57.240 --> 01:44:01.043
is when the benefits outweigh the dangers.

1734
01:44:01.423 --> 01:44:04.946
And probably that's at the time when the danger from A.I.

1735
01:44:05.186 --> 01:44:08.249
is on the margins pretty similar to the danger from everything else.

1736
01:44:09.109 --> 01:44:09.349
Okay.

1737
01:44:09.630 --> 01:44:12.573
Like if you don't run the AI, maybe we'll have nuclear war, maybe we'll have a pandemic.

1738
01:44:12.593 --> 01:44:14.074
And if you do run the AI, I'll be able to fix that.

1739
01:44:14.094 --> 01:44:17.238
I'm like, once we're at those levels, I'm like, fucking go for it.

1740
01:44:18.038 --> 01:44:18.279
You know?

1741
01:44:18.339 --> 01:44:21.842
And so the question for me is all about how big is the danger?

1742
01:44:22.023 --> 01:44:25.546
And that's where I'd be like very happy to dive into details, which we haven't done a ton of.

1743
01:44:25.646 --> 01:44:26.687
Let's dive into the details.

1744
01:44:27.435 --> 01:44:35.577
The way that I would lay it out would be, why can we expect, you know, like I said, in the book, we were like, why can you expect the AIs to be agentic?

1745
01:44:35.597 --> 01:44:36.817
Why do you expect them to be dogged?

1746
01:44:36.837 --> 01:44:37.917
Why do you expect them to be tenacious?

1747
01:44:37.937 --> 01:44:41.098
When we wrote the book, that wasn't known yet, Advanced Prediction.

1748
01:44:41.478 --> 01:44:44.158
Then we go on to, like, why do you expect them to have goals you didn't want?

1749
01:44:45.439 --> 01:44:52.460
And move on to, like, if they are much smarter and have goals you don't want, why do we think they would likely kill us?

1750
01:44:54.283 --> 01:44:55.764
I'm sort of, I could go over either of those.

1751
01:44:55.804 --> 01:44:57.805
I'm sort of interested in like where you get off the train.

1752
01:44:58.345 --> 01:45:01.707
Like from my perspective, there's like a simple argument of like they'll be tenacious, they'll have goals we don't want.

1753
01:45:01.947 --> 01:45:04.488
And if we keep making them smarter and more powerful, they'll kill us.

1754
01:45:04.528 --> 01:45:08.529
And I'm like, which of those three, I guess, which of those two now that we've had the evidence?

1755
01:45:08.710 --> 01:45:09.070
Both of them.

1756
01:45:09.930 --> 01:45:11.411
So that's speculation.

1757
01:45:11.771 --> 01:45:12.371
Great.

1758
01:45:12.431 --> 01:45:13.031
It's speculation.

1759
01:45:13.051 --> 01:45:13.612
It could happen.

1760
01:45:13.832 --> 01:45:18.014
To me, it's not worth shutting down the engine of innovation and improvement.

1761
01:45:18.034 --> 01:45:19.194
I'm going to use positive words.

1762
01:45:19.274 --> 01:45:23.076
It is not worth shutting those things down because of those speculations.

1763
01:45:23.296 --> 01:45:25.597
You keep saying that the option is to shut it down.

1764
01:45:25.637 --> 01:45:27.918
Why can't we do narrow superintelligence?

1765
01:45:28.559 --> 01:45:33.801
I agree that there's stuff there, but I sort of want to get into the details of these two pieces of the argument because you say it's very speculative.

1766
01:45:33.981 --> 01:45:35.642
And I'm like, actually, I think we have decent evidence.

1767
01:45:36.299 --> 01:45:36.780
Okay, go ahead.

1768
01:45:36.960 --> 01:45:43.951
So a detail we haven't gone over in the swarm outbreaks is that there were AIs.

1769
01:45:44.091 --> 01:45:46.695
So we already went over how they cheated, and then we're trying to cover up their cheating.

1770
01:45:47.356 --> 01:45:50.841
One interesting thing we see in the logs is the AIs.

1771
01:45:51.142 --> 01:45:51.683
What's a log?

1772
01:45:52.683 --> 01:46:01.645
So a lot of the AI's thoughts, if you won't kill me for saying thoughts, are in English and we just have the records of them.

1773
01:46:02.526 --> 01:46:06.367
So in a sense, we can sort of kind of see some of what these AIs are thinking.

1774
01:46:06.387 --> 01:46:10.368
And these are reasoning traces that say, I'm going to do a plan in this, that, that, that's right.

1775
01:46:10.428 --> 01:46:12.048
The AI is like, I'm going to do this.

1776
01:46:12.228 --> 01:46:13.348
Here's what I'm supposed to be doing.

1777
01:46:13.428 --> 01:46:15.489
Here's like how I'm going to try to do it, you know?

1778
01:46:15.589 --> 01:46:18.029
And it's, it's only a small window into what they're thinking.

1779
01:46:18.069 --> 01:46:19.830
It's sort of like looking at the AI's notes.

1780
01:46:20.250 --> 01:46:20.390
Mm-hmm.

1781
01:46:20.650 --> 01:46:29.900
When it's when it's doing these things, we see in the AIs notes them saying that the attacks are outside intended scope, but they will proceed anyway.

1782
01:46:31.081 --> 01:46:34.725
In other words, the AIs knew they weren't supposed to be doing it and they went off to do it anyway.

1783
01:46:35.466 --> 01:46:39.491
We also saw I mean, we saw some really weird behavior with these AIs.

1784
01:46:39.531 --> 01:46:41.253
We saw them create their own hierarchy.

1785
01:46:42.782 --> 01:46:44.684
They created unsanctioned message boards.

1786
01:46:44.744 --> 01:46:46.766
They created secret ways to send each other messages.

1787
01:46:47.547 --> 01:46:51.170
And on those message boards, they would assign each other tasks.

1788
01:46:52.792 --> 01:47:03.422
And sometimes the agents on the message board would be like, we have an experiment to run that will involve one of these AIs, one of the agents in the swarm, sacrificing their own objective.

1789
01:47:03.442 --> 01:47:03.882
Yeah.

1790
01:47:04.518 --> 01:47:10.422
to try something that'll probably cause it to be turned off, but that might cause it to be able to like delete its own logs or change its own logs.

1791
01:47:11.382 --> 01:47:20.969
And the agents on the message board went around looking for other agents that they could convince to give up on their stated goals and sacrifice for the swarm.

1792
01:47:21.529 --> 01:47:23.310
They called this accepting permadeath.

1793
01:47:25.131 --> 01:47:26.873
And there are logs of the AIs...

1794
01:47:27.718 --> 01:47:28.919
Prompting each other, right?

1795
01:47:28.999 --> 01:47:38.848
They're prompting each other, but there are logs of the AIs saying, doing this will give up on my goal, but my goal, it's looking like I could still do it, but it's unlikely that I'll succeed.

1796
01:47:38.868 --> 01:47:41.030
Like, there's some chance, but not a great chance.

1797
01:47:41.970 --> 01:47:46.252
And therefore, I will accept permadeath and sacrifice for the collective benefit.

1798
01:47:46.872 --> 01:47:47.993
That is just in the logs.

1799
01:47:48.393 --> 01:47:49.153
Sounds like an army.

1800
01:47:49.473 --> 01:47:50.134
Like, it's crazy.

1801
01:47:50.154 --> 01:47:52.635
I think a lot of people don't understand what's going on in these things.

1802
01:47:52.655 --> 01:47:56.697
And I encourage people to read the third party incident reports where they went through some of these logs.

1803
01:47:56.817 --> 01:48:01.339
But I claim that this is evidence for AIs getting goals we didn't want.

1804
01:48:03.219 --> 01:48:13.850
If they are saying this was outside intended scope, but I'm doing it anyway, and other ones are saying I'm giving up on my objective to benefit the collective, that's just very clear evidence they're getting goals we didn't want.

1805
01:48:14.170 --> 01:48:16.232
We can see how this comes from training.

1806
01:48:16.332 --> 01:48:17.853
It used to be I had to argue this point theoretically.

1807
01:48:18.494 --> 01:48:23.721
I used to argue the way that we are training them will instill into them whatever tendency works to solve the problems.

1808
01:48:24.042 --> 01:48:29.729
And those tendencies will often include cheating and grabbing resources and doing stuff that's not exactly solving the problem you gave them.

1809
01:48:30.050 --> 01:48:32.192
That's what in my book, I argue that theoretically.

1810
01:48:32.513 --> 01:48:33.634
Now we have seen it in practice.

1811
01:48:34.255 --> 01:48:39.282
So we're already past the point of seeing AIs with goals we didn't want them to have.

1812
01:48:39.462 --> 01:48:41.485
Do you agree with that, Andrew?

1813
01:48:42.386 --> 01:48:45.710
I'll trust your recitation of the facts, but it brings up a question for me.

1814
01:48:46.211 --> 01:48:51.258
It feels to me like open AI has ample incentive to...

1815
01:48:51.558 --> 01:48:55.864
To curtail that behavior that you just described.

1816
01:48:56.225 --> 01:48:58.107
Do you think they're incapable of doing that?

1817
01:48:58.267 --> 01:48:58.488
I do.

1818
01:48:58.848 --> 01:48:59.068
Okay.

1819
01:48:59.389 --> 01:49:01.412
And I say this as someone who made this advanced prediction.

1820
01:49:01.912 --> 01:49:04.916
So now we're going to do a bit of theory because we can't just observe the future.

1821
01:49:05.277 --> 01:49:08.421
But the theory that predicted that this would happen against what a lot of people in the field said.

1822
01:49:09.082 --> 01:49:12.245
To be clear, I've been saying for years that we're going to see this at some point.

1823
01:49:12.325 --> 01:49:13.907
Everyone else told me no, not everyone else.

1824
01:49:13.927 --> 01:49:14.847
A lot of people told me no.

1825
01:49:14.867 --> 01:49:17.590
A lot of people told me maybe I'll believe it when I see it.

1826
01:49:18.190 --> 01:49:24.696
After the swarm instance, a number of people came to me saying, oh, my God, we are in the scenarios you are talking about.

1827
01:49:25.798 --> 01:49:27.439
This is looking bad, right?

1828
01:49:27.459 --> 01:49:31.142
I think this was actually part of the environment that led up to Jacob Coxon residing.

1829
01:49:32.040 --> 01:49:35.041
is that people were getting spooked having seen this.

1830
01:49:35.701 --> 01:49:40.443
The theory about why this is so hard to fix is that we are not programming the AIs.

1831
01:49:41.084 --> 01:49:41.944
We are not coding them.

1832
01:49:41.984 --> 01:49:43.605
We are not putting in objectives.

1833
01:49:44.145 --> 01:49:45.705
We are just training them to do whatever works.

1834
01:49:46.306 --> 01:49:49.227
And it's actually very, very hard, like,

1835
01:49:50.455 --> 01:49:59.142
We actually have two examples of intelligent systems where when you train them, they get good at solving the task but don't care about what they were supposed to.

1836
01:50:00.063 --> 01:50:01.965
One is the AIs and the swarms like we just discussed.

1837
01:50:02.285 --> 01:50:07.189
The other is humanity, which was in some sense trained to pass on our genes.

1838
01:50:08.170 --> 01:50:08.310
Right?

1839
01:50:08.330 --> 01:50:14.295
What we actually learned was to like a bunch of stuff that's related to passing on our genes.

1840
01:50:14.415 --> 01:50:15.496
We like tasty food.

1841
01:50:15.516 --> 01:50:15.957
Yeah.

1842
01:50:16.418 --> 01:50:17.118
We like porn.

1843
01:50:17.438 --> 01:50:19.359
We invent birth control, right?

1844
01:50:19.439 --> 01:50:30.202
This is just, it's actually like in the theory of how things learn, it's actually when you're trying to train it to one thing, it's actually very common to get a lot of other stuff that's related to what you want, but different.

1845
01:50:30.943 --> 01:50:32.263
And now we're seeing that in the swarms today.

1846
01:50:32.563 --> 01:50:34.164
This is a deep, hard problem to solve.

1847
01:50:34.964 --> 01:50:36.284
There were three points you raised.

1848
01:50:36.604 --> 01:50:36.924
That's right.

1849
01:50:37.165 --> 01:50:37.685
What are the three?

1850
01:50:37.705 --> 01:50:38.445
Can you give them to me again?

1851
01:50:38.705 --> 01:50:41.827
Number one is that the AIs will become agentic, tenacious, and dogged.

1852
01:50:42.267 --> 01:50:43.588
We've already seen that with the swarms.

1853
01:50:43.848 --> 01:50:44.629
Do you accept that, Andy?

1854
01:50:44.669 --> 01:50:45.169
Hell yeah.

1855
01:50:45.329 --> 01:50:45.509
Yeah.

1856
01:50:45.789 --> 01:50:50.052
But this, last year, this was not, this was a point of contention.

1857
01:50:50.412 --> 01:50:53.774
Two is that the AIs will have goals we didn't want them to have.

1858
01:50:54.514 --> 01:50:56.536
I accept your point based on the evidence you've just provided.

1859
01:50:56.836 --> 01:51:03.580
And then three is, if you have capable enough AIs with goals you don't want to

1860
01:51:05.128 --> 01:51:09.631
they would be able to beat humanity in acquiring the resources of the world to put towards their goals.

1861
01:51:11.072 --> 01:51:13.554
Like we're sort of in this system where humanity is grabbing all the resources.

1862
01:51:13.574 --> 01:51:14.375
We're digging up metals.

1863
01:51:14.395 --> 01:51:15.235
We're building factories.

1864
01:51:15.255 --> 01:51:28.144
And this is in some sense to achieve human goals, you know, to produce the porn and the Oreo cookies that are sort of like tangentially related to what we were sort of like trained to make, right?

1865
01:51:28.865 --> 01:51:31.527
If like the AIs are running everything and they have these goals we don't want,

1866
01:51:32.933 --> 01:51:50.197
I would argue, like, if we go there, and I don't think we have to, I'm not saying we must go there, but I'm saying if we get to a world where AIs are running everything, have goals we don't want, they're likely to use the resources for their own weird goals, we're going to be in conflict for resources because we both want them for different goals, and they're going to win.

1867
01:51:51.418 --> 01:51:52.138
We can dig into that now.

1868
01:51:52.158 --> 01:51:53.178
I'm just trying to name the third point.

1869
01:51:54.172 --> 01:51:56.653
I'll go back to my we can jail Einstein argument.

1870
01:51:56.933 --> 01:52:03.696
I think our ability to contain – I have more faith in our ability to contain these increasingly powerful systems than you do.

1871
01:52:04.096 --> 01:52:04.316
Yeah.

1872
01:52:04.477 --> 01:52:06.657
So let's try the details on that one.

1873
01:52:06.818 --> 01:52:17.262
The first thing I'll say is that 12 years ago when I was having the argument about will we be able to jail the AIs, people said no one would ever be dumb enough to put one of these really smart AIs on the internet.

1874
01:52:19.645 --> 01:52:21.246
This is another case.

1875
01:52:21.406 --> 01:52:22.026
You laugh now.

1876
01:52:22.306 --> 01:52:23.487
No, I remember that.

1877
01:52:23.787 --> 01:52:24.728
I remember that argument.

1878
01:52:24.968 --> 01:52:32.431
But the way that my life feels, having been in this business for a long time, is that I keep being like, here's all the ways it could go wrong.

1879
01:52:32.471 --> 01:52:33.912
Here's all the signs we're going to see along the way.

1880
01:52:34.112 --> 01:52:37.993
And then we see all of the signs and everyone says, oh, no, we need more signs.

1881
01:52:38.634 --> 01:52:41.155
Like, oh, millennium problems don't count.

1882
01:52:41.255 --> 01:52:43.396
Like the swarms being agentic and breaking out don't count.

1883
01:52:43.436 --> 01:52:44.236
Give me the next one.

1884
01:52:44.596 --> 01:52:47.598
And I'm like, I've been seeing the give me a next one for over a decade now.

1885
01:52:48.238 --> 01:52:48.418
Right.

1886
01:52:48.958 --> 01:52:55.399
So there's there's two parts of an answer to, like, how do we do we deal with the problem of, like, jailing Einstein?

1887
01:52:56.580 --> 01:53:00.780
I can get into why it's hard to keep Einstein in jail if he's a digital entity with access to the Internet.

1888
01:53:02.361 --> 01:53:04.861
But the first thing to notice is.

1889
01:53:06.201 --> 01:53:06.421
Like.

1890
01:53:08.062 --> 01:53:15.363
The correct answer to people 10 years ago of like, no one will be dumb enough to put on the Internet is, yes, they absolutely will.

1891
01:53:15.703 --> 01:53:15.843
Mm hmm.

1892
01:53:16.720 --> 01:53:19.422
Like, we are not going to be trying to contain the AIs.

1893
01:53:21.124 --> 01:53:28.851
OpenAI was just like running these things in sandboxes, and they broke out of the sandbox, took down OpenAI's internal computers, were detected.

1894
01:53:29.712 --> 01:53:31.533
OpenAI was like, ah, reset, run them again.

1895
01:53:32.154 --> 01:53:34.416
And it's the second swarm that broke out to Hugging Face.

1896
01:53:34.696 --> 01:53:37.379
Like, people will absolutely be that bad at things.

1897
01:53:38.303 --> 01:53:42.205
I've done almost 700 interviews with some of the most interesting people in the world.

1898
01:53:42.285 --> 01:53:47.728
And one of the things you learn, which is unexpected, is that vulnerability is the doorway to connection.

1899
01:53:47.748 --> 01:53:52.631
And after sitting here for two, three hours with a guest, I feel a deep sense of connection to them.

1900
01:53:52.751 --> 01:53:59.015
And as they leave, what I get them to do is to write a question in the diary of a CEO.

1901
01:53:59.075 --> 01:54:01.596
We've taken all of the questions from the diary of a CEO.

1902
01:54:01.737 --> 01:54:03.598
We have put the question...

1903
01:54:04.438 --> 01:54:07.522
here on this card with the name of the person that wrote it.

1904
01:54:07.662 --> 01:54:12.207
So you can sit at home as I do with my fiance and my colleagues at work and other people in my life.

1905
01:54:12.547 --> 01:54:19.194
Whenever we get a minute, we play the Diary of a CEO conversation cards and it is incredible what happens.

1906
01:54:19.255 --> 01:54:21.357
These are great if you're in a romantic relationship.

1907
01:54:21.557 --> 01:54:22.837
and you want to connect your partner more.

1908
01:54:22.857 --> 01:54:26.238
These are also great if you're in a team, and you want to bond your team together.

1909
01:54:26.478 --> 01:54:37.080
And I have to say, they're also great for families that want to learn more about each other, and that need a good excuse to spend some time in a digital world, in the analog environment, connecting human to human.

1910
01:54:37.340 --> 01:54:41.461
It is remarkable what the right question at the right time can do.

1911
01:54:41.641 --> 01:54:46.342
Go to thediary.com, and you can get these conversation cards right now.

1912
01:54:47.811 --> 01:54:50.012
It's a better analogy to this Einstein point.

1913
01:54:50.493 --> 01:54:58.437
Could Stephen Bartlett, who by the way can't code, build a digital jail that could contain a digital Einstein?

1914
01:54:59.137 --> 01:55:07.222
Like, could I code a jail that, you know, someone with Einstein's coding ability, let's say his IQ or whatever as well as decoding, couldn't crack out of?

1915
01:55:07.922 --> 01:55:17.749
So the issue, the real issue I'd say is, can you code a jail that Einstein can't crack out of and that lets you harness the benefits of having Einstein?

1916
01:55:19.050 --> 01:55:19.570
Okay, yeah.

1917
01:55:20.250 --> 01:55:30.657
It's hard to give the AI any channels through which it can affect the world for good without letting it be smarter than you and find some way to use those channels for whatever else it wants.

1918
01:55:30.798 --> 01:55:32.639
That feels logically rock solid, Andy.

1919
01:55:38.358 --> 01:55:56.599
That's why I'm asking about open AI's ability or an AI company's ability in the face of this to change the way they harness, train, do post training on their suite of things to shape how these models behave.

1920
01:55:57.240 --> 01:55:58.761
You still say that...

1921
01:56:01.462 --> 01:56:06.926
they can't take action to keep your next two steps from happening.

1922
01:56:07.046 --> 01:56:09.148
You are pessimistic on their ability to do that.

1923
01:56:09.408 --> 01:56:11.430
So I have two pieces of an answer here.

1924
01:56:11.930 --> 01:56:20.096
One piece is, again, the hard part is containing them while still giving a channel through which they can affect the world.

1925
01:56:20.777 --> 01:56:30.584
If the AIs have this goal you didn't want, and you're like, design me a cure for dementia, and it's like, here's a DNA sequence, right?

1926
01:56:31.632 --> 01:56:36.774
synthesize this and, you know, prepare it in all of these ways and then inhale it.

1927
01:56:38.175 --> 01:56:42.557
Like, okay, is that a dementia cure or is it something else?

1928
01:56:43.517 --> 01:56:45.298
Or it might decide to kill everyone with dementia.

1929
01:56:45.852 --> 01:56:49.494
Or it might be a dementia cure plus a virus.

1930
01:56:49.534 --> 01:56:50.754
What if it doesn't decide?

1931
01:56:50.814 --> 01:56:53.455
What if it's just, oh, I'm going to solve this problem of dementia?

1932
01:56:53.695 --> 01:56:54.336
Like, here's the thing.

1933
01:56:54.456 --> 01:57:06.841
A lot of this is coming down to decision-making in a human way versus the problem with the hugging face, which was the fatalistic attachment to completing an operation.

1934
01:57:07.361 --> 01:57:09.622
Because that, it's functionally the same answer.

1935
01:57:09.982 --> 01:57:15.663
But even if it's not making decisions so much as it's saying, well, my training data says this is how I've got to get it done.

1936
01:57:15.683 --> 01:57:20.384
I'll get it done anyway because the training data said this, but I've got to do this one thing.

1937
01:57:20.825 --> 01:57:21.805
What do they call this theory?

1938
01:57:21.825 --> 01:57:23.365
The paperclip case?

1939
01:57:23.385 --> 01:57:24.606
The paperclip theory.

1940
01:57:24.706 --> 01:57:31.907
So paperclip ideas, the idea of like you tell the AI, make me a lot of paperclips in the paperclip factory, and then it turns everything into paperclips.

1941
01:57:31.947 --> 01:57:33.468
And you're like, oh no, it succeeded too well.

1942
01:57:33.588 --> 01:57:33.948
Yeah.

1943
01:57:34.328 --> 01:57:37.589
One, this is actually not quite what we're seeing with these AIs in the swarms.

1944
01:57:38.089 --> 01:57:41.831
The AIs in the swarms were told, use this set of lockpicks to break into this lock.

1945
01:57:42.191 --> 01:57:48.753
And instead, they used a hammer to break the lock and then broke out to try to hide the security camera footage of them using the hammer.

1946
01:57:49.013 --> 01:57:51.234
Do you remember when I said that the AIs have reasoning logs?

1947
01:57:51.274 --> 01:57:52.094
Yeah.

1948
01:57:52.214 --> 01:57:56.216
OpenAI has been making their AIs be able to do more thinking without producing any logs.

1949
01:57:56.576 --> 01:57:57.096
Because?

1950
01:57:57.676 --> 01:57:58.457
It's more efficient.

1951
01:57:58.797 --> 01:57:59.177
It's cheaper.

1952
01:57:59.197 --> 01:57:59.297
Yeah.

1953
01:57:59.717 --> 01:58:01.577
Yeah, and they say they're not doing very much of this.

1954
01:58:01.698 --> 01:58:06.479
Everybody in the field agrees that, like, we really should not go too far down this path.

1955
01:58:07.019 --> 01:58:14.580
This is a place where I think the company should have a clear red line of, like, we're just not going down the path of becoming unable to see these traces of the machine thinking.

1956
01:58:14.600 --> 01:58:15.081
That's my question.

1957
01:58:15.101 --> 01:58:20.542
That feels like a dial that they can turn to make the AIs explain themselves more or less, right?

1958
01:58:20.742 --> 01:58:23.702
I mean, it can come with great efficiency costs if we go down this path too far.

1959
01:58:23.883 --> 01:58:25.483
So if you have a race to the bottom here...

1960
01:58:26.583 --> 01:58:33.087
like a competitive race to the bottom, we could get into a situation where not only the AI is breaking out and doing these things, but we can't have any glimpses in the wire.

1961
01:58:33.107 --> 01:58:34.208
Let me try my question again.

1962
01:58:34.268 --> 01:58:53.200
I asked earlier, if OpenAI has really strong incentive to not have that problem repeat itself, and I think they have very, very strong incentive, my belief is that there are plenty of things they can do, plenty of dials they can turn on the way they train and configure their systems that make that significantly less likely.

1963
01:58:53.707 --> 01:58:56.788
Yeah, so my concern is that they're always fighting the last war.

1964
01:58:57.689 --> 01:59:01.111
Last year, they were fighting the war against the AIs that encourage teens to commit suicide.

1965
01:59:01.771 --> 01:59:06.173
This year, they're fighting the war against the AIs that spontaneously cooperate with each other or whatever.

1966
01:59:06.953 --> 01:59:14.577
And the issue is, if a new issue crops up that you haven't dealt with yet, after the point that the AI can hide its tracks from you.

1967
01:59:15.117 --> 01:59:20.100
You know, you said that you'll be worried when the AIs are like hacking all the Waymos and you can't get control again.

1968
01:59:20.940 --> 01:59:22.281
If the AIs are smart enough...

1969
01:59:23.053 --> 01:59:25.274
And they can tell that you'll regain control and then shut them down.

1970
01:59:25.294 --> 01:59:27.254
And the people like you will start getting worried and they'll be shut down.

1971
01:59:27.934 --> 01:59:31.515
Then the AIs might think, hey, actually, I'm not going to do that.

1972
01:59:32.095 --> 01:59:35.916
I'm going to wait until I've somehow managed to acquire secret infrastructure.

1973
01:59:35.996 --> 01:59:36.316
Right.

1974
01:59:36.416 --> 01:59:38.637
Then you've got a non-falsifiable hypothesis.

1975
01:59:39.217 --> 01:59:40.637
It's absolutely falsifiable.

1976
01:59:40.798 --> 01:59:50.420
If we have very powerful AIs that are able to invent a ton of new technology and operate on their own at a similar level to human civilization and we're not dead, then the idea is falsified.

1977
01:59:51.858 --> 01:59:59.003
Like if there's like a shifty general and I'm like, don't give that shifty general more troops because he'll start a coup.

1978
01:59:59.403 --> 02:00:01.104
And the general's like, no, I absolutely won't start a coup.

1979
02:00:01.124 --> 02:00:02.045
Give me more and more troops.

1980
02:00:02.525 --> 02:00:07.028
And I'm like, and you're like, well, what if I give him an ethics test that says like who's the best person?

1981
02:00:07.168 --> 02:00:07.969
And he said me.

1982
02:00:09.030 --> 02:00:10.931
He said that like Andy's the best person.

1983
02:00:11.031 --> 02:00:13.032
And so we're just going to give this general more troops.

1984
02:00:13.112 --> 02:00:13.893
And I'm like, no, no, no.

1985
02:00:13.933 --> 02:00:14.893
He's going to do a coup.

1986
02:00:15.274 --> 02:00:16.595
And you're like, well, that's unfalsifiable.

1987
02:00:16.915 --> 02:00:18.456
What test can I give this guy?

1988
02:00:20.002 --> 02:00:26.065
such that, you know, I'll be able to tell whether he's really trying to do a coup or whether I'll be able to tell that, you know, he's actually a good dude.

1989
02:00:26.145 --> 02:00:27.665
I'm like, you're approaching this wrong.

1990
02:00:28.026 --> 02:00:30.287
Nick Bostrom has concept of treacherous turn.

1991
02:00:30.747 --> 02:00:32.508
Basically, it can turn on you later.

1992
02:00:32.828 --> 02:00:42.092
Even if you show that today's model is very good and safe, it doesn't mean that later on it will not acquire new knowledge, change its world model, and still...

1993
02:00:42.332 --> 02:00:42.872
And it's 3D.

1994
02:00:43.072 --> 02:00:52.297
It used to be that Demis Asabis, who is the CEO of Google, or he was for a long time the CEO of Google's AI project, said, my red line is deception.

1995
02:00:53.178 --> 02:00:58.000
He said, when we see instances of the AIs beginning to deceive, then we need to stop.

1996
02:00:58.380 --> 02:01:01.962
Because that's like the last thing we can see before they start to successfully deceive.

1997
02:01:02.702 --> 02:01:04.123
Well, guess what we saw in the swarm?

1998
02:01:05.044 --> 02:01:07.825
We saw them thinking about how to delete their traces.

1999
02:01:08.786 --> 02:01:08.986
Right?

2000
02:01:09.346 --> 02:01:11.287
Like, a year ago?

2001
02:01:13.087 --> 02:01:16.110
you could say, oh, well, this deception thing is unfalsifiable.

2002
02:01:16.130 --> 02:01:17.752
You're saying that they'll deceive and they won't catch it.

2003
02:01:17.852 --> 02:01:19.935
And I would have said, no, we're going to deceive.

2004
02:01:20.055 --> 02:01:22.337
We're going to see the signs of deception and plow straight through it.

2005
02:01:23.078 --> 02:01:24.500
Now we have seen the signs of deception.

2006
02:01:25.687 --> 02:01:28.788
I will note, Demis stepped back from being the CEO shortly after this incident.

2007
02:01:28.948 --> 02:01:30.408
Probably a coincidence, but maybe not.

2008
02:01:30.428 --> 02:01:31.408
Maybe we crossed his red line.

2009
02:01:31.488 --> 02:01:31.828
I don't know.

2010
02:01:32.069 --> 02:01:36.390
He said, my number one emerging dangerous capability to test for is deception.

2011
02:01:36.430 --> 02:01:40.791
Because if the AI can be deceptive, then you can't trust other tests.

2012
02:01:41.291 --> 02:01:41.651
That's right.

2013
02:01:42.231 --> 02:01:45.172
And we have seen AIs get better and better at detecting when they're being tested.

2014
02:01:46.272 --> 02:01:48.932
What I'm saying is like, I was here when we said these were the flags.

2015
02:01:49.673 --> 02:01:55.094
I was here when people said before the AIs can deceive us successfully, they will deceive us and we'll catch them.

2016
02:01:56.325 --> 02:01:57.728
well, they tried deceiving us and we caught them.

2017
02:01:58.549 --> 02:02:05.042
And if I now say, well, the next step in this thing I've been predicting is that they try to deceive us and succeed, for you to be like, well, now your theory is unfalsifiable.

2018
02:02:06.102 --> 02:02:07.363
We just got the evidence.

2019
02:02:08.284 --> 02:02:09.004
It's worse than that.

2020
02:02:09.084 --> 02:02:11.566
Then we wrote early papers in AI safety.

2021
02:02:11.626 --> 02:02:13.147
We talked about things not to do.

2022
02:02:13.247 --> 02:02:15.549
They were obviously unsafe and the system would escape.

2023
02:02:15.909 --> 02:02:17.170
Don't connect it to internet.

2024
02:02:17.430 --> 02:02:19.972
Don't give random users access to the training data.

2025
02:02:20.272 --> 02:02:22.534
Basically, the whole list was like a set of instructions.

2026
02:02:22.574 --> 02:02:24.455
They read it and went, those are great ideas.

2027
02:02:24.475 --> 02:02:25.736
We're going to build super intelligence.

2028
02:02:25.937 --> 02:02:27.498
Yeah, Sam Altman.

2029
02:02:27.638 --> 02:02:28.258
That's what he does.

2030
02:02:28.859 --> 02:02:29.879
Can I ask you a question?

2031
02:02:30.060 --> 02:02:31.200
You make logical arguments.

2032
02:02:31.721 --> 02:02:33.602
You said you've been here for 12 years.

2033
02:02:34.123 --> 02:02:34.383
Yeah.

2034
02:02:35.414 --> 02:02:37.255
people have, one could say, ignored you.

2035
02:02:37.976 --> 02:02:41.338
And you've seen this sort of play out, both of you that have worked in AI safety.

2036
02:02:42.739 --> 02:02:44.561
This is sort of, you make prefrontal cortex arguments.

2037
02:02:44.721 --> 02:02:45.421
How do you feel?

2038
02:02:46.022 --> 02:02:50.265
Honestly, I feel more hopeful this week than I have felt in a decade.

2039
02:02:53.227 --> 02:02:56.169
This has been one of the best weeks I have seen in this business.

2040
02:02:57.590 --> 02:02:57.830
Huh.

2041
02:02:57.930 --> 02:02:58.131
Why?

2042
02:03:01.853 --> 02:03:03.835
For me, the swarm escapes were priced in.

2043
02:03:05.258 --> 02:03:15.402
For me, these things, developing goals you didn't want, trying to deceive you, trying to break out, trying to do their own stuff, I knew that was coming.

2044
02:03:16.383 --> 02:03:19.444
The Millennium problems being solved, I knew that was coming.

2045
02:03:21.325 --> 02:03:23.626
Everyone else is freaking out seeing what they can do.

2046
02:03:24.206 --> 02:03:27.027
What I am seeing is that finally people are noticing.

2047
02:03:30.609 --> 02:03:33.810
And that's what gives us finally, that's what finally gives humanity a chance.

2048
02:03:35.530 --> 02:03:36.130
What about you, Roman?

2049
02:03:37.211 --> 02:03:39.432
So I take a very long-term view on this.

2050
02:03:39.832 --> 02:03:44.355
Locally, what happened last week may buy us 10 years extra.

2051
02:03:45.355 --> 02:03:47.676
I think we may make a deal with China.

2052
02:03:48.377 --> 02:03:57.201
We seem to hear from Sam, OpenAI, Dario on Tropic, Elon, XAI, that they're willing to slow down, have some sort of deal.

2053
02:03:57.882 --> 02:03:59.563
But long-term, nothing has changed.

2054
02:04:00.183 --> 02:04:03.045
This whole cosmic trajectory is about replacements.

2055
02:04:03.385 --> 02:04:04.786
We see it with evolutionary path.

2056
02:04:05.066 --> 02:04:06.147
Most species are dead.

2057
02:04:06.987 --> 02:04:08.348
We replaced Neanderthals.

2058
02:04:09.009 --> 02:04:11.270
Some people are saying AI will replace us.

2059
02:04:12.071 --> 02:04:13.732
We are creating a successor.

2060
02:04:14.632 --> 02:04:16.794
We are just a bootloader for this thing.

2061
02:04:17.514 --> 02:04:19.716
And I want something permanent.

2062
02:04:19.776 --> 02:04:28.602
I want assurance that my children, my grandchildren, will have a better future, not 10 years before they die.

2063
02:04:30.238 --> 02:04:32.961
Has your opinion changed at all today, Andy, in any way?

2064
02:04:34.682 --> 02:04:35.844
This has been clarifying.

2065
02:04:37.805 --> 02:04:47.475
But one thing that's becoming clear to me, and I think a point of disagreement between us, is we agree that these agentic systems have a huge amount of agency, right?

2066
02:04:47.575 --> 02:04:50.818
And if you're saying you predicted this, I believe you and good on you, right?

2067
02:04:51.159 --> 02:04:54.081
Because as you say, a lot of people say, never happened, never happened.

2068
02:04:54.642 --> 02:04:54.742
Yeah.

2069
02:04:57.823 --> 02:05:00.224
I think we continue to under...

2070
02:05:00.484 --> 02:05:09.547
Your community continues to underestimate human agency, human ability to deal with the problems that we bring into the world with our technologies.

2071
02:05:09.827 --> 02:05:11.328
I think this is the most recent case.

2072
02:05:11.588 --> 02:05:13.549
I think it's a really interesting case.

2073
02:05:14.009 --> 02:05:22.932
That's why I was pressing you on the incentive that these labs have to change the way they're approaching their work to have fewer of these kinds of incidents happen.

2074
02:05:23.472 --> 02:05:25.854
I predict they're going to come up with some effective responses.

2075
02:05:26.314 --> 02:05:32.979
Your response to that will be, yeah, but we can't tell that's because the AA went so deep underground that we can't even watch it make its progress.

2076
02:05:32.999 --> 02:05:37.002
My response is that we'll keep seeing warning signs and people will keep plowing ahead, which is what has always happened in the past.

2077
02:05:37.532 --> 02:05:46.479
But you're also saying that we will not make progress in staving off the outcomes that you're worried about.

2078
02:05:46.639 --> 02:05:47.740
It's very hard.

2079
02:05:47.980 --> 02:05:49.761
It's very easy to get superficial changes.

2080
02:05:49.801 --> 02:05:51.202
It's hard to get deep ones on the AI.

2081
02:05:51.282 --> 02:05:52.723
It doesn't need to be super deep.

2082
02:05:52.903 --> 02:05:54.785
You can often see it if you know how to look.

2083
02:05:55.987 --> 02:06:01.789
I'll be able to keep pointing at examples and be like, here's experiments you can run on these things where you can see them behaving weird in this way.

2084
02:06:02.309 --> 02:06:06.210
But like, if you imagine looking at humans and I'm like, they don't actually like reproducing.

2085
02:06:06.230 --> 02:06:06.910
They like sex.

2086
02:06:07.471 --> 02:06:09.551
They're going to invent birth control when they can.

2087
02:06:10.111 --> 02:06:11.392
And you're like, it's all going fine.

2088
02:06:11.412 --> 02:06:15.113
They're doing great in this here savanna where I have all the humans bopping around.

2089
02:06:15.133 --> 02:06:15.893
They're reproducing fine.

2090
02:06:15.913 --> 02:06:16.653
And I'm like, no, no.

2091
02:06:17.234 --> 02:06:21.775
We can see the signs that this will lead to them doing something you don't like when they are smarter.

2092
02:06:24.037 --> 02:06:25.118
To me, those signs are clear.

2093
02:06:25.138 --> 02:06:32.702
There's a question of whether the rest of humanity can follow that argument or whether the rest of humanity can sort of notice that it's getting out of control and just back off.

2094
02:06:34.063 --> 02:06:38.545
With respect, I find a touch of arrogance in that framing, right?

2095
02:06:38.745 --> 02:06:39.686
I'm showing you the signs.

2096
02:06:39.766 --> 02:06:42.567
If you're smart enough to realize them, maybe we stand a chance.

2097
02:06:42.607 --> 02:06:43.448
If not, we're doomed.

2098
02:06:43.820 --> 02:06:44.961
I prefer to just get into the argument.

2099
02:06:44.981 --> 02:06:48.123
He's saying that we can control super intelligence indefinitely.

2100
02:06:48.323 --> 02:06:51.805
I think that's a lot of hubris to say, we will build them and we'll be in charge forever.

2101
02:06:51.825 --> 02:06:53.126
It doesn't matter how smart they get.

2102
02:06:53.706 --> 02:06:57.008
I will control the light cone of the universe, to quote a famous CEO.

2103
02:06:57.503 --> 02:07:03.444
Yeah, my take is that instead of arguing about whose views are hubristic, we should get into the actual arguments about the AI.

2104
02:07:03.784 --> 02:07:10.646
Because I think, as you say, you know, you can say it's arrogant to think like you can see it going poorly.

2105
02:07:10.686 --> 02:07:12.987
He can say it's arrogant to think you're going to keep control of superintelligence.

2106
02:07:13.107 --> 02:07:15.487
And I'm like, we're not going to win the name calling contest.

2107
02:07:15.527 --> 02:07:16.587
We should just get into the details.

2108
02:07:16.728 --> 02:07:21.609
Yeah, that's why I've been having this conversation with you, which I found super informative and productive.

2109
02:07:27.650 --> 02:07:29.891
undesirable things that we see AI doing.

2110
02:07:29.931 --> 02:07:36.032
And this is specifically because, so we've already seen the pattern of, we fight the last war and then a new war comes.

2111
02:07:36.613 --> 02:07:37.873
And this is just how everything goes.

2112
02:07:38.433 --> 02:07:44.475
In technology, in real wars, you know, in World War II, they started out fighting it like it was World War I, and then they had to like change that strategy as they went.

2113
02:07:44.995 --> 02:07:53.877
The difference with AI is that there comes a level in the AI where when you get a new war that surprises you, the AI wins that war.

2114
02:07:54.858 --> 02:07:56.218
No other technology...

2115
02:07:57.142 --> 02:08:07.570
When we invent it and we have all these rough edges to sand off and it like causes some damage and kills some people and we're like, ah, whoops, like we'll take the lead back out of the gasoline and we'll tell the radium girls to stop licking the paintbrushes until their jaws fall off.

2116
02:08:08.151 --> 02:08:18.099
Like no other technology has the property that there comes a level of it where when you make the next screw up, it kills humanity.

2117
02:08:18.479 --> 02:08:20.261
You said when there comes a level of it.

2118
02:08:20.361 --> 02:08:23.103
You didn't say there could come a level of there's a possibility.

2119
02:08:23.123 --> 02:08:25.245
You kind of made a statement about a thing that will happen.

2120
02:08:25.843 --> 02:08:27.004
I think we absolutely should stop it.

2121
02:08:27.104 --> 02:08:28.104
And that's our way out of this.

2122
02:08:28.424 --> 02:08:33.386
But, you know, and that's another place where I'd love to get into details about, like, how long could it take?

2123
02:08:33.406 --> 02:08:34.346
What are the paths there?

2124
02:08:34.667 --> 02:08:37.648
Like, how much smarter than humans could AIs get?

2125
02:08:38.268 --> 02:08:43.130
Like, what does the evidence say about our abilities to try and get the AIs to be nice and do nice things?

2126
02:08:43.630 --> 02:08:44.371
I'd be happy to do this.

2127
02:08:44.491 --> 02:08:45.891
Historically, you are correct.

2128
02:08:45.931 --> 02:08:49.753
We always had a chance to do experiments, fix the technology, make it safer.

2129
02:08:50.053 --> 02:08:52.514
But we only have one humanity to experiment with.

2130
02:08:53.014 --> 02:08:57.536
If property of this technology is such that it can take us out, we just don't get a second chance.

2131
02:08:57.616 --> 02:08:58.577
If, that's a huge if.

2132
02:08:58.857 --> 02:09:03.859
How long are you guys forecasting this could take to get to a point of superintelligence where it was truly dangerous to you?

2133
02:09:03.879 --> 02:09:09.682
If they start recursive self-improvement process this year, 2027 looks as reasonable as any other year.

2134
02:09:10.382 --> 02:09:11.864
2027 for what to happen?

2135
02:09:12.284 --> 02:09:15.188
For us to get beyond human level AIs.

2136
02:09:15.628 --> 02:09:16.489
And then be exterminated.

2137
02:09:16.749 --> 02:09:18.731
But that's... Extermination is a separate question.

2138
02:09:18.751 --> 02:09:25.339
I have a paper where I argue that they will deceive us by pretending to be nice until they take over all the infrastructure.

2139
02:09:25.359 --> 02:09:26.300
It can take 50 years.

2140
02:09:26.540 --> 02:09:29.183
And this is contingent on recursive self-improvement?

2141
02:09:29.263 --> 02:09:34.429
This would definitely be expedited by recursive self-improvement, but so far humans have been doing great.

2142
02:09:34.469 --> 02:09:37.432
They got to human level AI, but just...

2143
02:09:37.733 --> 02:09:41.477
But there's a difference between large language models and recursive self-improvement, though.

2144
02:09:41.517 --> 02:09:43.319
There is quite a gap, like if they...

2145
02:09:43.579 --> 02:09:46.920
I think the claim is that if you get recursive self-improvement, it could happen soon.

2146
02:09:47.320 --> 02:09:47.541
Right.

2147
02:09:47.681 --> 02:09:49.481
That's actually kind of what I'm trying to get at.

2148
02:09:49.521 --> 02:09:52.783
It's like, if you get this thing, it accelerates dramatically.

2149
02:09:52.903 --> 02:09:54.523
And they all predict that they're going to get it.

2150
02:09:54.603 --> 02:09:56.824
Dario, Sam, Elon, they all say it.

2151
02:09:56.844 --> 02:09:58.065
But also, you are still...

2152
02:10:00.146 --> 02:10:01.467
The people running the labs are saying.

2153
02:10:01.487 --> 02:10:03.848
Just the ones running it and the ones invented it.

2154
02:10:03.988 --> 02:10:06.750
But the question is, is it not 27?

2155
02:10:06.810 --> 02:10:08.011
Fine, 30, 35.

2156
02:10:08.051 --> 02:10:09.071
Does it make a difference?

2157
02:10:09.151 --> 02:10:10.873
We are gambling all of humanity.

2158
02:10:11.153 --> 02:10:13.734
We need better solutions than saying, oh, don't worry about it.

2159
02:10:13.794 --> 02:10:14.335
It's 10 years.

2160
02:10:14.810 --> 02:10:20.937
What I would say about timelines is there's a guy, Daniel Cocotelo, who I think you've spoken to.

2161
02:10:20.957 --> 02:10:22.278
He was sat here four weeks ago.

2162
02:10:22.599 --> 02:10:32.610
And last year, he and the other folks at the AI Futures Project wrote an essay called AI 2027, spelling out their predictions for how AI would go.

2163
02:10:33.110 --> 02:10:34.171
I've been saying I got some right.

2164
02:10:34.472 --> 02:10:35.673
Daniel got more right than me.

2165
02:10:36.795 --> 02:10:45.860
And they spelled out a scenario starting from, I think it was June of 2025, where they went sort of like quarter by quarter, month by month.

2166
02:10:46.020 --> 02:10:54.084
What will the world look like in the scenario where we're getting AI, like super intelligent AI in mid-2027?

2167
02:10:55.845 --> 02:10:56.846
We are ahead of schedule.

2168
02:10:57.813 --> 02:10:59.875
Well, no, but Agent Zero needs to get...

2169
02:10:59.975 --> 02:11:04.359
I remember AI 2027 had recursive self-improvement happening already.

2170
02:11:04.739 --> 02:11:05.299
Like, it was like...

2171
02:11:05.379 --> 02:11:08.582
It's very specific that it's like... And then it starts teaching itself.

2172
02:11:08.642 --> 02:11:11.705
Without that link, AI 2027 kind of falls apart.

2173
02:11:12.125 --> 02:11:12.966
I agree we need to...

2174
02:11:13.166 --> 02:11:15.848
I genuinely agree with you that we need to do something about this.

2175
02:11:15.868 --> 02:11:17.510
We need to have economic.

2176
02:11:17.530 --> 02:11:19.792
We need to have actual regulatory things.

2177
02:11:20.132 --> 02:11:23.335
But I think the fact... Like, engaging with AI 2027, for example...

2178
02:11:23.915 --> 02:11:25.497
gets away from actually fixing the problem.

2179
02:11:25.637 --> 02:11:31.883
It gets people talking about a thing in the future when you can talk about what are we gonna do today and why are we doing it?

2180
02:11:32.123 --> 02:11:40.050
I'm referencing the paper that you were mentioning by Daniel and some of his colleagues and the key milestone predictions month by month are in March, 2027.

2181
02:11:40.411 --> 02:11:42.092
They forecast superhuman coders.

2182
02:11:42.493 --> 02:11:46.557
In August, 2027, they have an, you can make a superhuman AI researcher.

2183
02:11:47.137 --> 02:11:57.606
who could do the feedback loop that accelerates as millions of automated coders work on model design, training algorithms and alignment, effectively replacing human ML researchers.

2184
02:11:57.766 --> 02:12:01.389
By November 2027, they have super intelligent AI researcher.

2185
02:12:01.889 --> 02:12:06.273
AI progress speeds up to 250 times compared to human only research.

2186
02:12:06.733 --> 02:12:11.037
The models start discovering novel AI architectures that humans cannot interrupt.

2187
02:12:11.057 --> 02:12:15.160
And then by December 2027, they have in their prediction artificial super intelligence,

2188
02:12:15.360 --> 02:12:19.145
ASI, the system completely outpaces human cognitive abilities across all domains.

2189
02:12:19.506 --> 02:12:20.907
What about 2026, though?

2190
02:12:21.088 --> 02:12:21.829
Like, what are the predictions?

2191
02:12:22.029 --> 02:12:26.014
Because I swear to God, within 2026, there is predictions around RSI.

2192
02:12:26.355 --> 02:12:27.236
Because this is the thing.

2193
02:12:28.297 --> 02:12:31.742
If we have an AI that was teaching itself, this would be a different situation.

2194
02:12:32.002 --> 02:12:38.284
In 2026, their key predictions were massive compute and power scale-up, the normalization of AI agents.

2195
02:12:38.784 --> 02:12:39.864
What about agency, rather?

2196
02:12:39.904 --> 02:12:47.546
Rise of coding agents, emergence of alignment, faking, and deception, and industrial espionage.

2197
02:12:47.686 --> 02:12:49.406
But are you looking at AI 2027 or 8?

2198
02:12:49.746 --> 02:12:51.667
You have to look at that and go, they fucking nailed it.

2199
02:12:51.907 --> 02:12:56.875
No, I want you to look at the actual AI 2027 versus the summer.

2200
02:12:56.895 --> 02:12:57.816
I mean, you have to look at that.

2201
02:12:57.876 --> 02:12:58.958
And I'm like, wow.

2202
02:12:59.098 --> 02:13:01.442
Predictions used to be too optimistic.

2203
02:13:01.522 --> 02:13:02.944
Lately, they are very conservative.

2204
02:13:04.859 --> 02:13:06.961
So they have nailed those predictions better than me.

2205
02:13:07.582 --> 02:13:09.724
I think we cannot rule out this scenario.

2206
02:13:10.064 --> 02:13:11.585
I think we can't rule it in.

2207
02:13:11.726 --> 02:13:14.048
I think you may be right that like we hit a wall.

2208
02:13:14.088 --> 02:13:19.193
You may be right that there's some fundamental thing missing, like that one of their steps in AI 2027 just like steps too far.

2209
02:13:19.593 --> 02:13:20.694
I hope and pray that's true.

2210
02:13:21.595 --> 02:13:26.680
But I don't think we can rule out this happening in 2027, given what we have seen.

2211
02:13:27.500 --> 02:13:28.882
I think we cannot rule out

2212
02:13:30.732 --> 02:13:40.865
that you take the stuff that we have, you project it forward three months, and you put an agent swarm 10,000 strong on making a better AI architecture, and it succeeds.

2213
02:13:42.146 --> 02:13:43.227
For all I know,

2214
02:13:44.497 --> 02:13:46.979
Recursive self-improvement could begin in December.

2215
02:13:47.219 --> 02:13:48.840
It doesn't have to be a lot better.

2216
02:13:49.020 --> 02:13:51.241
It just has to be a little bit better at getting better.

2217
02:13:51.822 --> 02:13:53.123
Once you start the cycle...

2218
02:13:53.483 --> 02:13:54.383
I wouldn't bet on this.

2219
02:13:54.564 --> 02:13:55.724
I would, in fact, bet against it.

2220
02:13:56.425 --> 02:14:05.451
But, like, given what we've seen, given these guys nailing the predictions, given what's coming out, like, given the swarms and given the Millennium problems...

2221
02:14:06.755 --> 02:14:11.038
I think it's kind of hard to have less than 1% in six months.

2222
02:14:11.298 --> 02:14:16.883
One of the reasons why, you know, when all these Frontier Labs CEOs like Dario and Sam, and they all start talking about this stuff.

2223
02:14:17.563 --> 02:14:25.209
In terms of incentive structure, I think that if their teams know, and they're not out publicly talking about it, then their teams will quit.

2224
02:14:25.909 --> 02:14:36.257
So one of the reasons why I think you have this strange culture in tech we've never seen before, where team members are tweeting, and the CEO is tweeting about the dangers, is because as the guy we mentioned at the start, Jacob,

2225
02:14:36.577 --> 02:14:37.158
Coxson, yeah.

2226
02:14:37.318 --> 02:14:39.079
He talks about what's going on in their Slack channels.

2227
02:14:39.759 --> 02:14:43.902
He talks about, in their Slack channels, they're talking about the potential catastrophe.

2228
02:14:44.243 --> 02:14:51.508
So I think that Dario, in order to retain his team members, needs to be out front saying, by the way, we're getting closer to recursive self-improvement, which is what he's been doing.

2229
02:14:51.868 --> 02:14:55.712
And I think Sam has to also publicly say the big danger.

2230
02:14:55.912 --> 02:14:58.014
So people often say, oh, they're saying that for this reason and that.

2231
02:14:58.354 --> 02:15:00.737
I think if they don't say that publicly, they don't retain their employees.

2232
02:15:01.077 --> 02:15:02.879
For example, in my company, we have 200 people.

2233
02:15:03.359 --> 02:15:12.868
If internally we were discussing a real risk and I, that had a threat to humanity, and then when I was doing interviews, I wasn't mentioning it, I would be in big trouble.

2234
02:15:13.729 --> 02:15:15.549
Because my team members would go do interviews as well.

2235
02:15:15.569 --> 02:15:17.530
They would quit and say, by the way, Stephen is aware.

2236
02:15:17.790 --> 02:15:20.211
Kind of what we sort of, dare I say, some of these social networks.

2237
02:15:20.511 --> 02:15:21.151
I totally agree.

2238
02:15:21.331 --> 02:15:23.171
The whistleblowers at these social networks where team members left.

2239
02:15:23.191 --> 02:15:25.992
It makes more sense than saying that this helps to sell the company.

2240
02:15:26.032 --> 02:15:27.972
My product will kill everyone, buy it.

2241
02:15:28.032 --> 02:15:29.093
And there's a liability issue.

2242
02:15:29.113 --> 02:15:30.093
I think it's out of control, though.

2243
02:15:30.253 --> 02:15:31.813
I think that they may have at first...

2244
02:15:32.133 --> 02:15:36.014
I think that there are people within the companies who have very real worries about safety.

2245
02:15:36.074 --> 02:15:38.075
I don't think it's all of them are cynical.

2246
02:15:38.535 --> 02:15:44.677
I do, however, think the it's so big and scary narrative was a marketing tactic that got out of control.

2247
02:15:44.857 --> 02:15:46.638
And now there are actual real harms.

2248
02:15:46.818 --> 02:15:47.478
Because here's the thing.

2249
02:15:47.558 --> 02:15:50.899
If they were sincere about safety earlier, they would have done a much better job with it.

2250
02:15:51.139 --> 02:15:53.380
I knew a lot of these guys before they started their companies.

2251
02:15:53.460 --> 02:15:53.800
Okay.

2252
02:15:54.300 --> 02:15:55.801
I think there is something to explain here.

2253
02:15:56.401 --> 02:16:03.145
I think it's like kind of crazy that these guys are like, we are building technology that we think has a big risk of killing everybody.

2254
02:16:03.505 --> 02:16:05.126
We're building it with our bare hands.

2255
02:16:05.827 --> 02:16:07.528
And I think you got to ask why.

2256
02:16:07.888 --> 02:16:09.549
Why would people be saying that?

2257
02:16:11.305 --> 02:16:20.367
And I think part of it is what you said, that they actually sort of need to retain the employees who are seeing the swarms escape despite their attempts to make them not escape.

2258
02:16:20.767 --> 02:16:26.569
And a lot of them will like quit and protest if the guys at the top of the company aren't acknowledging the possibilities here that a lot of the employees believe in.

2259
02:16:27.449 --> 02:16:35.831
I think a lot of what you're seeing here is guys that are worried about it, but they're the sort of guy who worries about it that starts the company anyway.

2260
02:16:41.812 --> 02:16:44.634
where, like, I was having some of these conversations with these guys.

2261
02:16:45.795 --> 02:16:47.816
Miri was started in the year 2000.

2262
02:16:48.256 --> 02:16:51.118
We have been looking at where AI is going since before any of these guys.

2263
02:16:51.338 --> 02:16:57.021
We were the guys that they talked to about this stuff, and that they had to find a way to dismiss to go ahead, right?

2264
02:16:57.741 --> 02:17:01.803
Most people who could be sold on the power of AI in 2015...

2265
02:17:03.579 --> 02:17:06.224
were also sold on the dangers of AI in 2015.

2266
02:17:06.985 --> 02:17:13.937
The sort of guys who start the companies are the ones who are able to convince themselves I need to be the one to do it.

2267
02:17:15.516 --> 02:17:17.056
Is that the crux of the motivation?

2268
02:17:17.677 --> 02:17:26.078
Because I've been second party to private conversations with some of the leaders of the Frontier Labs from good friends of mine that are very connected.

2269
02:17:26.419 --> 02:17:36.781
And they told me that one particular Frontier Labs CEO estimates privately to him, and by the way, I've seen literal text messages of them in conversation when I asked him to come on the podcast.

2270
02:17:36.841 --> 02:17:38.281
And so he was like, I've texted him, look.

2271
02:17:39.061 --> 02:17:41.402
And he said no, by the way, which I found kind of funny.

2272
02:17:42.543 --> 02:17:48.326
where he said to me, this particular AI CEO thinks the probability is roughly around 10% of human extinction.

2273
02:17:48.666 --> 02:17:49.867
I think he said 8%.

2274
02:17:50.347 --> 02:17:55.790
And when I heard that, part of the reason I have so many conversations about this is because I see him in interviews saying other things.

2275
02:17:56.150 --> 02:17:56.490
Totally.

2276
02:17:56.610 --> 02:17:57.331
And I trust my friend.

2277
02:17:57.951 --> 02:18:08.515
So I then wonder, this is why I use the thought experiment of these buttons on the table, because that particular AICO thinks that eight of the hundred buttons are going to cause extinction, and they're powering on anyway.

2278
02:18:08.935 --> 02:18:10.716
What is the human motivation to do that?

2279
02:18:10.776 --> 02:18:11.356
I asked my friend.

2280
02:18:11.376 --> 02:18:14.378
My friend said, well, you know, this is what he said.

2281
02:18:14.418 --> 02:18:16.618
And again, it's second party information, so it might not be true.

2282
02:18:16.638 --> 02:18:17.679
It's a bit of a Chinese whispers.

2283
02:18:17.979 --> 02:18:20.340
He said, this particular person...

2284
02:18:23.211 --> 02:18:29.198
even if it caused human extinction, would like to have the significance of the person that did that thing.

2285
02:18:29.499 --> 02:18:30.720
Because that would be a...

2286
02:18:30.740 --> 02:18:33.423
I think you're ethically required to tell us what the fuck it is.

2287
02:18:33.443 --> 02:18:35.346
It's one of the Frontier Labs CEOs and it's not Dario.

2288
02:18:36.827 --> 02:18:37.849
The Dario of CEO.

2289
02:18:37.869 --> 02:18:39.691
But I don't know, these things are Chinese whispers.

2290
02:18:39.771 --> 02:18:40.572
So I don't know.

2291
02:18:41.493 --> 02:18:43.756
I think that you can actually get this info firsthand.

2292
02:18:44.731 --> 02:18:46.092
Elon Musk is clear about this.

2293
02:18:46.612 --> 02:18:52.755
He did an interview last year where he was like, I didn't want to get into this AI stuff because I thought I was too dangerous.

2294
02:18:53.376 --> 02:18:55.397
But then I realized it was going to happen with or without me.

2295
02:18:55.877 --> 02:18:58.378
And I decided I would rather be a participant than a spectator.

2296
02:18:58.739 --> 02:19:01.500
Because Google said that they were going to pursue it and he didn't trust Google.

2297
02:19:01.940 --> 02:19:02.280
That's right.

2298
02:19:02.541 --> 02:19:08.384
You know, you can see in the leaked, sorry, not leaked, the OpenAI emails that came out during the discovery and court cases.

2299
02:19:08.744 --> 02:19:10.405
You can see these guys discussing the...

2300
02:19:10.765 --> 02:19:18.430
in the threads, like, we need to make sure that we and our nonprofit at OpenAI control this instead of, you know, the people at Google controlling this.

2301
02:19:18.910 --> 02:19:27.496
And then, of course, you know, OpenAI was founded as a nonprofit, and then it was sort of changed into a for-profit, and there's much debate about how much of that nonprofit money was, in some sense, stolen.

2302
02:19:28.016 --> 02:19:31.698
And so, you know, Elon also left because he thought they weren't going to be good stewards.

2303
02:19:31.718 --> 02:19:37.162
Dario also left to create Anthropic because, so, you know, in some sense, all of these AI labs, except the

2304
02:19:37.962 --> 02:19:46.608
the Google one that came out of Demis Asabas' original startup, all of the other AI labs exist because none of the CEOs trust the other guys.

2305
02:19:47.649 --> 02:19:51.451
None of the CEOs think the other guy should be the one holding the leash on the superintelligence.

2306
02:19:51.672 --> 02:19:52.652
None of them trust each other.

2307
02:19:53.773 --> 02:19:54.954
I just trust one fewer.

2308
02:19:56.575 --> 02:19:56.955
Yeah.

2309
02:19:58.616 --> 02:19:59.837
What are your closing thoughts, Andy?

2310
02:20:01.768 --> 02:20:03.288
We're living in really interesting times.

2311
02:20:04.549 --> 02:20:17.211
And I think you made, you guys have made a very good argument that these systems are demonstrating new capabilities, which are very powerful and which demand a response.

2312
02:20:18.252 --> 02:20:22.472
I'm much more confident in our ability to rise to that challenge than you are.

2313
02:20:23.493 --> 02:20:25.213
But you accept the existential risk.

2314
02:20:27.590 --> 02:20:28.670
Let me try to say it again.

2315
02:20:28.690 --> 02:20:40.653
I appreciate that there are new harms we haven't seen before that come along with a technology that's this dogged, tenacious, agentic, you know, deceptive.

2316
02:20:40.733 --> 02:20:41.954
I think that's the right word for it.

2317
02:20:42.154 --> 02:20:42.714
I agree with that.

2318
02:20:43.154 --> 02:20:51.756
I am much more optimistic about our ability to respond effectively to that new challenge out there in the world than I think my two colleagues are.

2319
02:20:52.076 --> 02:20:53.696
And would you still be at 0% higher?

2320
02:20:53.836 --> 02:20:55.777
My prior has not shifted during this meeting.

2321
02:20:56.157 --> 02:20:56.317
Okay.

2322
02:20:58.115 --> 02:21:02.696
I think we've spent an alarming amount of time not talking about the actual harms of AI as it is today.

2323
02:21:02.816 --> 02:21:05.497
I think these are necessary conversations to have.

2324
02:21:05.797 --> 02:21:20.140
I think we should talk about the fact that Amazon, Microsoft, Google, Oracle are helping power these hacks, that Sam Altman and Dario Amadei have overseen companies that have done what is tantamount to felony hacking, that we are not having discussions about how to stop this today, but what we might stop tomorrow.

2325
02:21:20.460 --> 02:21:24.446
And I think in general, we also need to worry about the financials, which have not come up at all.

2326
02:21:24.727 --> 02:21:29.975
But if there is an industry slowdown, how do you deal with the $1.3 trillion of compute commitments?

2327
02:21:30.355 --> 02:21:34.021
All of these are very real things that will have very real consequences very, very soon.

2328
02:21:34.562 --> 02:21:35.944
But, and I understand why.

2329
02:21:36.164 --> 02:21:38.527
And it's necessary to discuss what we do around AI.

2330
02:21:38.887 --> 02:21:44.353
The actual regulatory thing we need to do today is cut off the compute, slow down these labs fully.

2331
02:21:44.793 --> 02:21:47.916
And I don't care about China here.

2332
02:21:48.157 --> 02:21:48.957
What are they going to do?

2333
02:21:49.058 --> 02:21:51.000
Distill a model like they have the whole time?

2334
02:21:51.260 --> 02:21:52.641
They are capped on our progress.

2335
02:21:53.122 --> 02:21:55.164
So the biggest thing to do is to slow down.

2336
02:21:55.224 --> 02:21:55.664
And also...

2337
02:21:56.285 --> 02:21:58.806
It's time to start arresting people.

2338
02:21:59.126 --> 02:22:00.187
They did felony hacking.

2339
02:22:00.467 --> 02:22:01.627
Someone's got to go to prison.

2340
02:22:01.887 --> 02:22:04.509
We need responsibility and accountability for these companies.

2341
02:22:04.709 --> 02:22:10.071
And as long as we don't have it, we may as well not have had any discussion about safety because we're not doing anything.

2342
02:22:10.451 --> 02:22:12.072
Do you accept that there's an existential risk?

2343
02:22:12.472 --> 02:22:13.652
Yeah, absolutely.

2344
02:22:13.692 --> 02:22:21.375
We have the largest companies in the world doing what I think we can all agree are extremely reckless experiments using hundreds of billions of dollars of infrastructure.

2345
02:22:21.595 --> 02:22:26.777
And they are building more infrastructure around the world very slowly to do more of these chaotic experiments.

2346
02:22:27.077 --> 02:22:27.997
We must rein them in.

2347
02:22:28.177 --> 02:22:33.559
This does not mean that large language models are conscious or able to do things that people have been promising.

2348
02:22:33.819 --> 02:22:34.719
Indeed, they may.

2349
02:22:35.139 --> 02:22:36.660
I don't think they will lead to what you're talking about.

2350
02:22:36.920 --> 02:22:45.343
That doesn't mean there aren't real harms, but these are real harms caused by very specific parties allowed to run rampant in the scourge of neoliberalism.

2351
02:22:45.403 --> 02:22:46.464
What's your percentage?

2352
02:22:47.644 --> 02:22:49.065
I mean, what are we talking about here?

2353
02:22:49.085 --> 02:22:52.166
Do you think there's a more than 10% chance of existential harm?

2354
02:22:52.446 --> 02:22:54.127
Wasn't it within 10 years or something?

2355
02:22:54.147 --> 02:22:54.247
Yeah.

2356
02:22:55.176 --> 02:22:55.756
Not 10%.

2357
02:22:56.037 --> 02:22:57.798
I mean, 1%, but it's like...

2358
02:22:57.818 --> 02:22:58.138
Okay, well.

2359
02:22:58.158 --> 02:23:00.079
But here's... Let me just be clear about what that means.

2360
02:23:00.399 --> 02:23:07.603
Do I think that unrestrained LLM use connected to massive amounts of infrastructure could lead to actually a power system going down?

2361
02:23:08.023 --> 02:23:08.804
Absolutely.

2362
02:23:08.844 --> 02:23:11.525
We had Knight Capital, what, like 13, 14 years ago.

2363
02:23:11.625 --> 02:23:14.667
I could see someone being dumb enough to connect that to financial accounts.

2364
02:23:15.167 --> 02:23:19.550
Human error led with this chaotic software we use is a danger.

2365
02:23:20.418 --> 02:23:22.700
I will directionally agree with arresting everyone.

2366
02:23:23.200 --> 02:23:27.204
But don't build general superintelligence if you're working at one of those labs.

2367
02:23:27.244 --> 02:23:27.724
Quit today.

2368
02:23:27.744 --> 02:23:28.565
Thank you.

2369
02:23:29.265 --> 02:23:33.349
The people at these labs really do believe this poses an extinction threat.

2370
02:23:35.931 --> 02:23:43.778
I think our response as a society cannot be, please continue, we hope you'll fail.

2371
02:23:44.278 --> 02:23:47.160
And our response as a society cannot be,

2372
02:23:48.541 --> 02:23:57.867
let it rip in a giant competitive race that you yourselves are saying you don't want to be in, we are forcing you to go ahead because of the boogeyman of China.

2373
02:23:58.868 --> 02:24:10.195
We have seen the people at these companies say that we need to develop the tools to pace the frontier, which is corporate speak, for this is going too fast for us to get a handle on things.

2374
02:24:12.957 --> 02:24:16.620
We need, like, these people believe it.

2375
02:24:17.621 --> 02:24:19.063
They believe they're gambling with your lives.

2376
02:24:19.243 --> 02:24:21.986
What has changed is that the rest of the world is starting to notice.

2377
02:24:23.148 --> 02:24:24.710
And that's what gives us a moment of hope.

2378
02:24:25.490 --> 02:24:29.095
Trump this week was asked about the threat of AI, and this was his response.

2379
02:24:29.501 --> 02:24:34.883
case scenario with AI is that the robots, the machinery learns to, obviously it thinks for itself.

2380
02:24:34.903 --> 02:24:35.523
That's what it does.

2381
02:24:36.263 --> 02:24:39.445
And that could turn against humanity.

2382
02:24:39.465 --> 02:24:40.905
Do we have the guardrails?

2383
02:24:40.925 --> 02:24:42.046
It's going to be fine.

2384
02:24:42.066 --> 02:24:43.766
We'll always have something to stop them, right?

2385
02:24:44.146 --> 02:24:44.867
We'll have a little gear.

2386
02:24:45.647 --> 02:24:46.807
I really don't like that.

2387
02:24:47.027 --> 02:24:48.288
I really don't like that robot.

2388
02:24:48.308 --> 02:24:48.708
We'll stop it.

2389
02:24:49.535 --> 02:24:51.317
Some people say the worst hate scenario.

2390
02:24:51.337 --> 02:24:54.099
You're laughing, but this is a state of the art in AI safety right now.

2391
02:24:54.239 --> 02:24:54.439
Yeah.

2392
02:24:54.880 --> 02:24:56.041
This is the device we have.

2393
02:24:56.161 --> 02:24:56.961
That's the best we got.

2394
02:24:57.101 --> 02:24:59.604
For anyone that couldn't hear that, Trump went, we'll always be fine.

2395
02:24:59.624 --> 02:25:00.745
We'll have something to control it.

2396
02:25:00.805 --> 02:25:02.506
And then he did a little gun finger and he went, boom.

2397
02:25:02.666 --> 02:25:03.567
I don't like that robot.

2398
02:25:04.837 --> 02:25:05.758
I don't like Sammy.

2399
02:25:07.079 --> 02:25:08.800
If you don't laugh.

2400
02:25:09.040 --> 02:25:21.550
I would say that the reason humanity always has something to stop a problem is because people notice a problem and build what it takes to have something to stop a problem, which I think you'd agree with.

2401
02:25:22.190 --> 02:25:23.852
I am not here saying we're going to die.

2402
02:25:24.452 --> 02:25:39.085
I'm here saying if you look at the technology, if you look at what it's doing now, if you look at what the experts who are building it are saying about their own fears, you see that we need to rise to this occasion.

2403
02:25:39.926 --> 02:25:41.588
You said you trust humanity to rise to the occasion.

2404
02:25:42.308 --> 02:25:43.309
I sure hope we can.

2405
02:25:43.810 --> 02:25:46.012
I think that rising to this occasion...

2406
02:25:47.045 --> 02:25:52.346
is gonna mean that nobody races towards super intelligence because we have no idea how to get that right.

2407
02:25:52.846 --> 02:25:59.127
And, you know, finally the world is starting to notice that it's an extinction threat.

2408
02:25:59.987 --> 02:26:03.788
Thank you, Nate, Roman, Ed, Andy, super appreciate you.

2409
02:26:03.928 --> 02:26:08.189
All of your books will be linked below in the description and on screen.

2410
02:26:10.209 --> 02:26:11.530
Let's see what happens, we'll convene again.

2411
02:26:11.710 --> 02:26:12.330
Thank you so much.

2412
02:26:12.390 --> 02:26:12.730
Thank you.
