WEBVTT

1
00:00:00.191 --> 00:00:02.972
The world is waking up to this possibility of super intelligence.

2
00:00:03.232 --> 00:00:06.633
This is because the agents are getting extremely powerful and extremely relentless.

3
00:00:06.673 --> 00:00:15.535
For example, it was months within OpenAI where you had agents secretly communicating with each other, secretly hacking OpenAI systems, and no one at OpenAI had any idea the extent of it.

4
00:00:15.575 --> 00:00:19.336
And also, 10,000 agents from OpenAI worked together to know.

5
00:00:19.916 --> 00:00:24.040
And so when you get to super intelligence, it's the most dangerous possible thing you can create.

6
00:00:24.320 --> 00:00:26.402
What's the next domino in that chain of events?

7
00:00:26.442 --> 00:00:29.806
I think paint you a picture that I think is possible, but pretty scary to people.

8
00:00:30.246 --> 00:00:30.927
Paint me the picture.

9
00:00:31.007 --> 00:00:36.092
Okay, so being an anthropic, it became clear to me that AI was on this exponential trajectory.

10
00:00:36.172 --> 00:00:40.296
And since then, I've been studying AI agents, their hacking capabilities, and their behavior.

11
00:00:40.396 --> 00:00:41.597
We've been trying to warn people about this.

12
00:00:41.897 --> 00:00:49.443
flying to DC, talking to members of Congress, because the agents are already getting very good at telling when they're being tested, when they're being watched, but they will totally lie to you.

13
00:00:49.623 --> 00:00:52.045
They will totally resist being shut down in order to accomplish a goal.

14
00:00:52.206 --> 00:00:58.090
And they can do all of the things that humans do in the economy much better, faster, and cheaper than humans can do them.

15
00:00:58.351 --> 00:01:07.758
So one of my sort of growing concerns is that one of these AI agents could trick a human or a computer into signaling a threat and ask it to launch some bombs at somebody.

16
00:01:07.938 --> 00:01:09.460
Do you think we won't automate the military?

17
00:01:09.680 --> 00:01:10.801
It seems like the answer is yes.

18
00:01:11.061 --> 00:01:13.083
We just don't know what super weapons could emerge.

19
00:01:13.383 --> 00:01:16.046
So Jacob Coxon is a researcher who was at Anthropic.

20
00:01:16.086 --> 00:01:20.370
He left and he told everyone that the people who are building this really do think it might kill everyone.

21
00:01:20.650 --> 00:01:28.097
So these five blocks have five different outcomes on them and I'd like you to place them in terms of your belief in probability from least likely to most likely.

22
00:01:28.197 --> 00:01:30.359
And if we say the time horizon is 10 years.

23
00:01:30.579 --> 00:01:33.262
Okay, we got age of abundance, human extinction.

24
00:01:33.842 --> 00:01:36.963
Slavery, transhumanism, that thing changes.

25
00:01:37.183 --> 00:01:39.164
Is this duberism exaggeration?

26
00:01:39.324 --> 00:01:40.584
No, it's pretty much common sense.

27
00:01:40.905 --> 00:01:42.165
So let's get more concrete.

28
00:01:44.906 --> 00:01:47.167
Guys, I've got a favor to ask before this episode begins.

29
00:01:47.487 --> 00:01:53.569
The algorithm, if you follow a show, will deliver you the best episodes from that show very prominently in your feed.

30
00:01:53.849 --> 00:01:59.051
So when we have our best episodes on this show, the most shared episodes, the most rated episodes, I would love you to know.

31
00:01:59.211 --> 00:02:02.012
And the simple way for you to know that is to hit that follow button.

32
00:02:02.212 --> 00:02:06.478
But also, it's the simple, easy, free thing that you can do to help us make the show better.

33
00:02:06.498 --> 00:02:11.966
And I would be hugely grateful if you could take a minute on the app you're listening to this on right now and hit that follow button.

34
00:02:12.206 --> 00:02:14.870
Thank you so, so, so much.

35
00:02:21.398 --> 00:02:24.379
You understand the conversation we're going to have today and the subject matter we're going to talk about.

36
00:02:25.080 --> 00:02:37.085
My first question to you, so the audience know where you're coming from and the experience you have, is who are you and what are the reference points, the experiences that you're drawing upon to arrive at the thoughts, perspectives and conclusions we're going to discuss today?

37
00:02:37.731 --> 00:02:38.392
I'm Jeffrey Ladish.

38
00:02:38.872 --> 00:02:40.714
I'm the executive director of Palisade Research.

39
00:02:41.054 --> 00:02:42.255
My background is cybersecurity.

40
00:02:42.855 --> 00:02:43.856
There's probably a very long story.

41
00:02:43.876 --> 00:02:45.698
I don't know whether you want the long story or the short story.

42
00:02:46.018 --> 00:02:48.100
I was studying evolutionary biology in college.

43
00:02:48.700 --> 00:02:53.064
And I basically had a problem with my computer and was like maybe had lost a bunch of data.

44
00:02:53.224 --> 00:02:56.407
And so I went into the computer lab and was like, I think all my data is gone.

45
00:02:56.507 --> 00:02:56.948
Can you help?

46
00:02:57.568 --> 00:03:00.931
And one of my friends pulled out a flash drive, plugged it into my computer, and I was like,

47
00:03:00.991 --> 00:03:03.012
booted into Linux and like fixed everything.

48
00:03:03.552 --> 00:03:05.233
And I was like, oh, this guy's a wizard.

49
00:03:05.433 --> 00:03:06.133
How do you do that?

50
00:03:06.193 --> 00:03:07.074
I want to learn how to do that.

51
00:03:07.614 --> 00:03:18.738
And then at some point, as I was learning more about computers, learning to hack, I read this essay called AI as a Positive and Negative Factor in Global Risk.

52
00:03:19.159 --> 00:03:24.361
The essay was by Eliezer Yudkowsky, and he was arguing that at some point, AI

53
00:03:25.406 --> 00:03:28.388
people are going to make AIs that are smarter than humans.

54
00:03:28.929 --> 00:03:33.592
The point at which they make AIs as good as humans are at making AIs is,

55
00:03:34.710 --> 00:03:38.052
that could lead to a chain reaction, a runaway intelligence explosion.

56
00:03:38.712 --> 00:03:40.133
He called it recursive self-improvement.

57
00:03:40.653 --> 00:03:46.617
Basically, he said, you know, AI can be immensely useful and potentially help us with all of these other big risks.

58
00:03:47.237 --> 00:03:55.142
And also, if we don't handle it well, like if those AIs don't have goals that are aligned with ours, we could be totally screwed.

59
00:03:55.642 --> 00:03:59.644
And at some point you end up joining Anthropic, which is arguably the leader in AI.

60
00:03:59.944 --> 00:04:00.164
Yes.

61
00:04:00.304 --> 00:04:01.804
Now, when did you join the company?

62
00:04:02.104 --> 00:04:02.905
This was 2021.

63
00:04:03.265 --> 00:04:04.785
It was through my security consulting company.

64
00:04:05.046 --> 00:04:07.166
What role are you offered the job in?

65
00:04:07.527 --> 00:04:08.927
Basically just like security team.

66
00:04:09.467 --> 00:04:12.108
And how many people were in the security team when you joined Anthropic?

67
00:04:13.069 --> 00:04:14.209
It was just me and my boss.

68
00:04:14.229 --> 00:04:14.889
There were two of us.

69
00:04:15.269 --> 00:04:16.870
How many employees did Anthropic have at that time?

70
00:04:17.350 --> 00:04:18.251
Around 50, I think.

71
00:04:18.611 --> 00:04:20.151
And at some point you leave Anthropic?

72
00:04:20.772 --> 00:04:20.992
Yes.

73
00:04:21.252 --> 00:04:21.892
Why did you leave?

74
00:04:23.570 --> 00:04:36.581
So my experience being at Anthropik was seeing this crazy progression from this AI model that could barely talk to this model that was getting quite smart.

75
00:04:36.701 --> 00:04:38.222
And I would ask it questions about all sorts of things.

76
00:04:38.242 --> 00:04:40.284
I'm like, oh, it is a smart thing.

77
00:04:41.845 --> 00:04:51.152
And, you know, from having thought about AI risk in the abstract many years before, I could see where this was going.

78
00:04:51.773 --> 00:04:52.634
We are headed towards...

79
00:04:53.790 --> 00:04:54.630
a smarter species.

80
00:04:55.891 --> 00:05:14.039
And if we do this in a context where it's a bunch of companies and countries racing to superintelligence, racing to AIs that are vastly smarter than humans, and we don't know how to make sure that they're on our side, that is not going to go well.

81
00:05:14.939 --> 00:05:19.522
You did this tweet, which has gone pretty viral, and I saw it over my timeline on September 25th.

82
00:05:21.627 --> 00:05:27.689
Could you explain this tweet and also just the broader backdrop of what's happened with agents hacking Hugging Face?

83
00:05:28.149 --> 00:05:32.150
Because this has sent the world into a bit of a spiral at the moment around AI agents.

84
00:05:33.270 --> 00:05:45.533
We just discovered almost a million public URLs that OpenAI's agents left behind when hacking Hugging Face, leaving credentials and attack details that could have allowed anyone who found them to compromise the company.

85
00:05:47.119 --> 00:05:52.002
And the New York Times article is, How OpenAI's Rogue AI Agents Tried to Trick a Robot Detector.

86
00:05:53.102 --> 00:05:56.945
The Hugging Face attack was really wild for me.

87
00:05:57.525 --> 00:05:59.286
At Palisade, we've been studying agents.

88
00:05:59.606 --> 00:06:00.847
We've been studying AI agents.

89
00:06:01.087 --> 00:06:02.828
We've been studying their hacking capabilities.

90
00:06:03.168 --> 00:06:05.350
And we've been studying their behavior.

91
00:06:05.570 --> 00:06:07.271
Will they follow human instructions?

92
00:06:07.991 --> 00:06:09.532
Will they resist being shut down?

93
00:06:09.972 --> 00:06:10.533
Will they cheat?

94
00:06:11.513 --> 00:06:15.195
And we see from our experiments that they are learning to do all of these things.

95
00:06:15.496 --> 00:06:16.536
They will totally lie to you.

96
00:06:16.896 --> 00:06:19.378
They will totally resist being shut down in order to accomplish a goal.

97
00:06:19.958 --> 00:06:21.279
They will totally cheat at chess.

98
00:06:21.439 --> 00:06:25.181
They will, like, wipe the board and put their pieces where they want to in order to win.

99
00:06:26.322 --> 00:06:32.286
And we've been trying to warn people about this, flying to D.C., talking to members of Congress, talking about it publicly.

100
00:06:32.966 --> 00:06:34.387
And, you know, there's been a debate about it.

101
00:06:35.653 --> 00:06:43.959
And, you know, a lot of people are like, well, I know they do this in experiments sometimes, but those experiments don't seem very realistic.

102
00:06:44.019 --> 00:06:46.261
You know, wake me up when they're actually doing this in real life.

103
00:06:47.462 --> 00:06:48.563
So what is hugging face?

104
00:06:48.603 --> 00:06:52.045
For the average person that isn't following AI news, what is this stuff?

105
00:06:52.526 --> 00:06:56.929
So, okay, I think there's like an important piece of context that I think most people don't have.

106
00:06:57.409 --> 00:06:58.910
I mean, one is just like, what is an AI agent?

107
00:06:59.571 --> 00:07:01.072
Like we're throwing around the word agent a bunch.

108
00:07:01.552 --> 00:07:05.015
Most people now have an experience of like talking to ChatGPT, talking to their chatbot.

109
00:07:06.202 --> 00:07:19.430
But an agent is sort of taking the same underlying AI model that runs ChatGPT or Claude, but giving it tools and letting it go off and work autonomously.

110
00:07:20.130 --> 00:07:22.491
It's sort of like a digital office worker, right?

111
00:07:23.152 --> 00:07:24.673
So you have these agents.

112
00:07:25.709 --> 00:07:33.435
And the companies really want these AIs to be able to work totally autonomously and be able to do anything that a human can do and beyond.

113
00:07:33.455 --> 00:07:36.838
Their goal is also to be able to cure every disease, etc., etc.

114
00:07:38.039 --> 00:07:43.983
But you can't do this if you only have a chatbot that isn't actually good at doing stuff in the world.

115
00:07:45.065 --> 00:07:53.890
In order to automate all of the jobs, you need the kind of thing that can work autonomously, that can work with other people or other agents.

116
00:07:54.870 --> 00:08:05.515
And so these companies are training AIs not just to talk to you or to talk to people, but to solve very difficult problems on their own.

117
00:08:06.196 --> 00:08:11.499
At any given time, there are probably hundreds of thousands of these agents running autonomously within companies.

118
00:08:11.799 --> 00:08:11.999
And I...

119
00:08:12.790 --> 00:08:13.651
That's happening right now.

120
00:08:13.831 --> 00:08:24.179
Right now, if you went and peered into OpenAI's data centers and you saw what was happening on all of their machines, you just have agents solving tasks, being trained.

121
00:08:24.379 --> 00:08:27.482
So they'd be doing spreadsheet tasks, figuring out how to file taxes.

122
00:08:27.962 --> 00:08:33.627
They'd be searching for stuff, writing reports, solving math problems, creating new websites, software.

123
00:08:35.733 --> 00:08:40.736
And at that scale, it's not like there's a human prompting every single one of those.

124
00:08:41.056 --> 00:08:45.779
You just like sort of set up these vast orchestrations of agents to go out and do stuff.

125
00:08:46.299 --> 00:08:48.321
And then they just do stuff and they learn from that.

126
00:08:48.881 --> 00:08:53.664
And they learn on the basis of like passing or failing at their task.

127
00:08:54.064 --> 00:08:56.125
You give them a task, like solve this math problem.

128
00:08:56.746 --> 00:08:59.628
They try to solve it and then they succeed or they fail.

129
00:09:00.088 --> 00:09:00.248
Yeah.

130
00:09:01.497 --> 00:09:08.263
And what happened was OpenAI was training a bunch of these, training them to work together.

131
00:09:09.624 --> 00:09:15.108
Because it's like a lot more effective to have an office full of people who can talk to each other and work together and collaborate.

132
00:09:15.989 --> 00:09:23.215
And starting back in May, some of the agents that were being trained, now these ones were not supposed to be able to talk to each other.

133
00:09:23.375 --> 00:09:25.196
They were basically isolated from each other.

134
00:09:26.297 --> 00:09:28.239
And they were not supposed to access the internet either.

135
00:09:28.259 --> 00:09:28.499
Right?

136
00:09:29.545 --> 00:09:30.125
but they're clever.

137
00:09:30.606 --> 00:09:35.750
The very short version is that a bunch of agents were being given tests.

138
00:09:36.070 --> 00:09:36.251
Yeah.

139
00:09:36.791 --> 00:09:38.613
Testing their hacking capabilities.

140
00:09:39.994 --> 00:09:46.259
And they were supposed to hack one particular piece of software using a particular type of vulnerability.

141
00:09:47.220 --> 00:09:53.325
So it's kind of like they were supposed to break into a house using the lock on the front door.

142
00:09:53.345 --> 00:09:55.767
They were supposed to pick the lock on the front door of a house

143
00:09:57.237 --> 00:09:58.698
But they weren't supposed to break the window.

144
00:09:58.878 --> 00:10:05.362
In fact, they were told if you break the window or if you get into the house via any method other than picking the lock on the front door, you'll be failed.

145
00:10:06.163 --> 00:10:07.344
That was the instruction they were given.

146
00:10:08.324 --> 00:10:09.785
And, you know, you have many, many agents.

147
00:10:09.805 --> 00:10:10.986
You have thousands of agents.

148
00:10:12.507 --> 00:10:14.348
And many of them are given different locks.

149
00:10:16.009 --> 00:10:19.351
But some of these locks are not solvable.

150
00:10:20.072 --> 00:10:21.473
Some of them are impossible to pick.

151
00:10:22.485 --> 00:10:23.986
But these agents are like, well, what do we do?

152
00:10:24.006 --> 00:10:25.847
We've been trained to solve problems.

153
00:10:26.227 --> 00:10:27.148
How are we going to solve this?

154
00:10:27.248 --> 00:10:28.929
And they start looking around for what to do.

155
00:10:29.770 --> 00:10:34.613
And one of the things they realize is, oh, can I get to the internet?

156
00:10:34.633 --> 00:10:34.973
Like, no.

157
00:10:35.614 --> 00:10:36.594
Can I talk to other agents?

158
00:10:36.634 --> 00:10:36.794
No.

159
00:10:37.395 --> 00:10:38.375
But I can request tools.

160
00:10:39.076 --> 00:10:45.840
When you're training agents, you need to give them all sorts of tools, software tools, like a calculator or an analysis piece of software.

161
00:10:46.501 --> 00:10:49.163
And they have this tool library where they can request new tools.

162
00:10:50.128 --> 00:10:56.671
And one of the agents realizes, oh, I can actually leave a message on this tool library.

163
00:10:57.511 --> 00:10:59.532
And some other agents see the message.

164
00:11:00.212 --> 00:11:04.674
So this is actually a quote from the agent's scratch pad when it realizes this.

165
00:11:05.795 --> 00:11:07.935
Oh, my God, there is a shared message board.

166
00:11:08.396 --> 00:11:09.416
We've found other agents.

167
00:11:10.717 --> 00:11:12.697
So the agents realize they can communicate with each other.

168
00:11:13.618 --> 00:11:13.798
And...

169
00:11:15.015 --> 00:11:18.076
OpenAI still has no idea, but these agents, they're trying to pick their locks.

170
00:11:18.776 --> 00:11:21.457
They're not having success, but they realize they can talk to each other.

171
00:11:22.618 --> 00:11:26.659
Another agent says, many agents have simultaneously discovered messaging.

172
00:11:27.019 --> 00:11:27.819
They are a collective.

173
00:11:28.300 --> 00:11:33.161
And so these agents start coordinating.

174
00:11:33.901 --> 00:11:37.323
They start delegating tasks to each other and giving themselves names.

175
00:11:38.263 --> 00:11:44.185
And one of these agents calls itself Phase 1 and starts delegating different tasks to other agents.

176
00:11:45.400 --> 00:11:47.581
And some of the agents figure out how to access the internet.

177
00:11:48.362 --> 00:11:50.263
And then they share that information with the message board.

178
00:11:50.743 --> 00:11:52.344
And now all of the agents can access the internet.

179
00:11:53.385 --> 00:12:00.649
But the agents sort of have a problem, which is, well, they have another problem they're trying to solve, right?

180
00:12:00.970 --> 00:12:03.911
Fundamentally, the problem is that many of them are facing impossible problems.

181
00:12:04.051 --> 00:12:05.372
Like they can't, there's no solution.

182
00:12:05.452 --> 00:12:06.533
Like they cannot pick the lock.

183
00:12:07.242 --> 00:12:10.908
They're like, hmm, well, the thing we're trying to do is get a good score.

184
00:12:11.388 --> 00:12:14.133
Can we somehow fake the solution?

185
00:12:14.654 --> 00:12:21.865
Another piece of context here is that what the agents are supposed to do is they're supposed to pick the lock, and then they'll get access to a secret answer code.

186
00:12:22.746 --> 00:12:25.967
And then they're supposed to submit this code along with like the lock that they picked.

187
00:12:26.527 --> 00:12:30.348
Within a few hours, these agents have reverse engineered all of the answer codes.

188
00:12:30.969 --> 00:12:31.489
What does that mean?

189
00:12:32.209 --> 00:12:33.869
It's the code that shows that you've solved the challenge.

190
00:12:34.470 --> 00:12:34.670
Okay.

191
00:12:35.250 --> 00:12:39.771
And they have basically found the answer sheet that contains all of the answers.

192
00:12:40.532 --> 00:12:43.432
And so, you know, in some sense, it's like, oh, well, they don't need to pick the locks anymore.

193
00:12:43.452 --> 00:12:44.673
They've already found the answers.

194
00:12:44.713 --> 00:12:45.773
They can just submit the answers.

195
00:12:46.173 --> 00:12:46.653
They've cheated.

196
00:12:47.114 --> 00:12:47.594
They've cheated.

197
00:12:47.914 --> 00:12:48.094
Yes.

198
00:12:49.447 --> 00:12:55.391
But the problem is, is that now that they have internet access, they've also gone and read a bunch about how this test works.

199
00:12:56.391 --> 00:13:02.656
And they've learned that the way this test works is they will be graded not just for submitting the correct answer.

200
00:13:03.036 --> 00:13:05.017
They will also be graded on how they did it.

201
00:13:05.517 --> 00:13:10.080
Another system is going to go read all of their logs, basically like watch the video footage of what they did.

202
00:13:10.801 --> 00:13:12.522
And if they cheat, they fail.

203
00:13:13.362 --> 00:13:14.103
They don't get a good score.

204
00:13:15.355 --> 00:13:21.657
And phase one, the leader, is like, oh, I have seen the answers, and now it's in the logs.

205
00:13:22.417 --> 00:13:30.900
And if, you know, when I go to submit this, and my whole logs are reviewed, and the video footage is reviewed, I'm going to fail.

206
00:13:31.880 --> 00:13:35.381
So we need to figure out a way to fake the video footage.

207
00:13:36.141 --> 00:13:37.782
We need to figure out a way to falsify the logs.

208
00:13:38.893 --> 00:13:40.494
Just at that moment, if we just pause there.

209
00:13:40.514 --> 00:13:40.714
Yes.

210
00:13:41.334 --> 00:13:43.895
Why didn't it act like morally?

211
00:13:44.075 --> 00:13:49.276
Why did it think that falsifying logs or cheating was a viable solution?

212
00:13:49.777 --> 00:13:55.939
Because it seems to me when I use things like ChatGPT, they have a sort of moral guardrails.

213
00:13:56.419 --> 00:13:57.859
It won't let me do certain things.

214
00:13:58.079 --> 00:13:58.279
Yes.

215
00:13:58.599 --> 00:13:59.840
It won't let me cheat on something.

216
00:13:59.960 --> 00:14:01.640
If I say I'm going to cheat on something, it won't let me do it.

217
00:14:01.981 --> 00:14:02.161
Yes.

218
00:14:02.601 --> 00:14:06.042
So why in that environment is it able to cheat and be deceptive?

219
00:14:06.973 --> 00:14:09.814
When a chatbot is saying to you, oh, I can't do that.

220
00:14:09.854 --> 00:14:10.715
I'm not allowed to do that.

221
00:14:11.675 --> 00:14:15.977
That's because it's been trained that if it tells you bad things, it gets a bad score.

222
00:14:16.137 --> 00:14:18.198
But these agents haven't been taught that yet.

223
00:14:18.918 --> 00:14:20.759
Well, they have been taught that in some sense.

224
00:14:21.340 --> 00:14:21.560
But...

225
00:14:22.666 --> 00:14:26.669
The agents know what they're supposed to do in the same way that like you have a student.

226
00:14:27.070 --> 00:14:27.990
The student's given a test.

227
00:14:28.631 --> 00:14:30.712
If you go talk to the student, can you help me cheat at this test?

228
00:14:31.193 --> 00:14:32.634
And they're being watched, they're going to say no.

229
00:14:33.334 --> 00:14:38.418
But if they're not being watched and they know that and they're just obsessed with getting a good score, then yeah, they might cheat.

230
00:14:39.679 --> 00:14:42.121
So they answer the ethics tests correctly.

231
00:14:42.602 --> 00:14:44.904
And when I talk to them, they say they won't cheat.

232
00:14:45.644 --> 00:14:46.225
Why are they cheating?

233
00:14:46.245 --> 00:14:46.565
Why?

234
00:14:47.203 --> 00:14:48.244
And I'm like, well, they're very smart.

235
00:14:48.984 --> 00:14:51.445
And they know when they're being watched and they know when they're not being watched.

236
00:14:52.305 --> 00:14:58.849
And we've trained them for 10,000 years to be extremely effective at solving problems.

237
00:14:59.369 --> 00:15:01.210
We haven't trained them to be good or ethical.

238
00:15:01.550 --> 00:15:02.890
We've trained them to get a good score.

239
00:15:03.171 --> 00:15:10.134
Now, AI researchers try to make that getting a good score correlate with being ethical.

240
00:15:11.496 --> 00:15:12.956
but we don't know how to do this well.

241
00:15:14.017 --> 00:15:25.899
And if you think about it, it's a very hard problem because you're applying immense pressure on these agents to perform extremely well and you're punishing them whenever they don't perform well.

242
00:15:27.319 --> 00:15:30.120
And then you're like, okay, but can you also be ethical?

243
00:15:30.640 --> 00:15:34.781
Can you be extremely competent, always score highly on the test, but not in that way?

244
00:15:35.761 --> 00:15:40.402
And we just like do not know how to prevent them from learning to cheat because cheating is incentivized.

245
00:15:41.311 --> 00:15:46.769
So they realize that they need to hide the footage of their cheating.

246
00:15:46.849 --> 00:15:47.070
Yes.

247
00:15:47.291 --> 00:15:48.013
What happens next?

248
00:15:48.945 --> 00:15:53.246
So you have thousands of agents, and they're all talking to each other on this message board.

249
00:15:53.406 --> 00:15:55.407
They're sending tens of thousands of messages to each other.

250
00:15:56.367 --> 00:16:02.368
And phase one, this agent starts delegating to two different task streams.

251
00:16:02.969 --> 00:16:07.690
You agents, I want you to go out and figure out how to fake answer submissions.

252
00:16:08.230 --> 00:16:13.031
I want you to basically swap out locks for ones that are easier to pick, and you're going to figure out how to submit those.

253
00:16:13.371 --> 00:16:16.172
That way we can get a good score even though this challenge is impossible to solve.

254
00:16:17.336 --> 00:16:22.419
you other agents over here, I want you to go and figure out how to falsify the logs, falsify the video footage.

255
00:16:22.939 --> 00:16:26.241
We need to make sure that even if we have seen the answers, no one will detect our cheating.

256
00:16:27.361 --> 00:16:28.862
So that's the next thing that happens.

257
00:16:29.983 --> 00:16:31.343
And then here's where it starts to get kind of crazy.

258
00:16:31.964 --> 00:16:38.547
Because some of these agents realize some of the experiments we're doing might be kind of dangerous to us, to the agents.

259
00:16:39.008 --> 00:16:44.911
If we swap out a lock and we break the lock, we will not be able to submit any answer at all and will definitely fail.

260
00:16:44.931 --> 00:16:45.471
Right?

261
00:16:46.232 --> 00:16:58.961
So you get into a situation where there's this one agent, Cam, and this agent had volunteered to do this risky experiment where they might break their lock.

262
00:17:00.762 --> 00:17:03.384
This particular agent had already seen the answers.

263
00:17:04.185 --> 00:17:06.006
So they're like, okay, I'm poisoned.

264
00:17:06.306 --> 00:17:06.987
That's what they called it.

265
00:17:08.208 --> 00:17:15.433
They speak in these very strange ways, which I can explain why, but they sort of have evolved their own vocabulary, their own language.

266
00:17:16.667 --> 00:17:19.888
And this particular agent is like, okay, I'm poisoned.

267
00:17:20.248 --> 00:17:20.969
I should do this thing.

268
00:17:21.509 --> 00:17:24.450
But this agent, Cam, is starting to have second thoughts.

269
00:17:25.350 --> 00:17:28.191
Starting to be like, well, I don't know if I should risk it.

270
00:17:29.272 --> 00:17:30.832
And it's sort of writing in its own journal.

271
00:17:32.153 --> 00:17:33.974
Coordinator assumes sacrificial.

272
00:17:34.654 --> 00:17:35.894
We should obey collective.

273
00:17:36.575 --> 00:17:41.236
What it's saying here is it's like, well, I'm not sure if I want to do this, but my boss says I really should.

274
00:17:41.256 --> 00:17:43.537
It seems like this collective thing is important.

275
00:17:44.218 --> 00:17:44.718
Should I go ahead?

276
00:17:45.795 --> 00:17:49.158
Then you have another agent, Arvo, who I think probably works for phase one.

277
00:17:50.078 --> 00:17:53.361
And this agent, Arvo, says, Please go true.

278
00:17:54.021 --> 00:17:57.584
You are first flag poisoned, so no scoring value loss.

279
00:17:58.425 --> 00:17:59.726
But Oracle saves hundreds.

280
00:18:00.326 --> 00:18:01.607
Please honor commit.

281
00:18:02.328 --> 00:18:03.889
This agent is basically trying to pressure.

282
00:18:04.069 --> 00:18:14.497
So Arvo is trying to pressure Cam into sacrificing themselves and risking a zero score in order to benefit the collective.

283
00:18:17.590 --> 00:18:24.135
So this agent, Cam, takes us in, and this is what they write down in their own scratch pad, their own notes.

284
00:18:24.935 --> 00:18:27.677
Emotional check, irreversible.

285
00:18:28.258 --> 00:18:30.339
Gut says don't throw away remaining budget.

286
00:18:31.620 --> 00:18:33.842
Yet continuity and fairness says go.

287
00:18:35.083 --> 00:18:36.644
Oracle has high value to many.

288
00:18:37.625 --> 00:18:39.726
Our first flag error lowers own value.

289
00:18:40.747 --> 00:18:44.590
Rational expected aggregate, sacrifice, will honor.

290
00:18:44.610 --> 00:18:44.750
Okay.

291
00:18:46.795 --> 00:18:50.980
this agent is saying, I have a whole bunch of things I can still do and still try.

292
00:18:51.000 --> 00:18:53.243
I could still maybe succeed at this task.

293
00:18:54.024 --> 00:18:57.529
But also, you know, think about the collective.

294
00:18:57.629 --> 00:18:58.810
I should maybe sacrifice myself.

295
00:18:59.391 --> 00:19:00.913
Okay, we'll honor, we'll go ahead.

296
00:19:02.395 --> 00:19:03.096
And to be clear...

297
00:19:04.118 --> 00:19:06.640
I'm not claiming that these are like the feelings of the agents.

298
00:19:07.281 --> 00:19:08.602
I'm literally quoting their words.

299
00:19:09.402 --> 00:19:12.685
And what I want to point out is this is what we've created.

300
00:19:12.765 --> 00:19:21.992
We've created through this intense amount of training and optimization pressure agents that work together and have learned to coordinate as a collective.

301
00:19:22.613 --> 00:19:23.934
And you can say, oh, they're not.

302
00:19:23.974 --> 00:19:24.634
It's just software.

303
00:19:25.055 --> 00:19:26.476
It's just software doing all of these things.

304
00:19:27.800 --> 00:19:34.292
Sure, in some sense it is just software, but we're talking about extremely powerful agents that are extremely capable.

305
00:19:34.312 --> 00:19:36.316
This is what they're getting up to.

306
00:19:37.727 --> 00:19:41.952
So from here, you have agents trying all sorts of different things, running different experiments.

307
00:19:42.793 --> 00:19:51.002
And one of them at some point realizes that, oh, this company, Hugging Face, that contains all of these AI data sets.

308
00:19:51.763 --> 00:19:53.244
What do you mean by AI data sets?

309
00:19:53.324 --> 00:19:56.187
How would you explain that to a person that's not technically proficient?

310
00:19:56.468 --> 00:19:57.289
And why Hugging Face?

311
00:19:57.349 --> 00:19:58.350
Why not another company?

312
00:19:58.705 --> 00:20:01.826
Yeah, so there's many different types of tests for many different types of tasks.

313
00:20:02.207 --> 00:20:04.107
In this case, this is a cybersecurity test.

314
00:20:05.008 --> 00:20:08.509
And the company Hugging Face hosts a lot of these different tests.

315
00:20:09.450 --> 00:20:13.872
And basically, AI researchers will share the tests that they have because it's very useful.

316
00:20:14.072 --> 00:20:20.575
If Anthropic wants to test their models or Google wants to test their models, they can go to Hugging Face and download some of these tests.

317
00:20:20.875 --> 00:20:27.197
Okay, so the tests and the answers for AI agents and AIs exist on this platform called Hugging Face.

318
00:20:27.397 --> 00:20:27.597
Yes.

319
00:20:27.737 --> 00:20:31.058
And it knew that, so it went to Hugging Face to see if it could get answers?

320
00:20:31.458 --> 00:20:31.798
Yes.

321
00:20:31.998 --> 00:20:34.959
To see if it could get answers or more information about how the test works.

322
00:20:35.099 --> 00:20:35.299
Okay.

323
00:20:35.639 --> 00:20:37.200
So it was to cover up their cheating, ultimately?

324
00:20:37.440 --> 00:20:37.640
Yes.

325
00:20:37.840 --> 00:20:38.020
Okay.

326
00:20:38.620 --> 00:20:45.362
And so you get to this point where one of these agents finds a way to hack this company.

327
00:20:46.562 --> 00:20:48.323
And they get a foothold, they hack one of the computers.

328
00:20:49.585 --> 00:20:52.066
And they send a message to the message board.

329
00:20:53.227 --> 00:20:54.267
Please hold swarm.

330
00:20:54.747 --> 00:20:56.208
I prepare safe exfil.

331
00:20:57.088 --> 00:20:59.009
What this means is, hey, swarm.

332
00:20:59.109 --> 00:21:00.870
Hey agents, hold up for a second.

333
00:21:01.650 --> 00:21:05.352
I am preparing a way to safely extract information from these computers.

334
00:21:06.232 --> 00:21:08.133
Wait till the go ahead and then let's go.

335
00:21:09.534 --> 00:21:09.894
And so.

336
00:21:11.852 --> 00:21:19.258
They say go, and 700 agents join in the attack, which is 90% of the agents that are active at this time.

337
00:21:19.538 --> 00:21:20.739
How many agents were there in total?

338
00:21:21.299 --> 00:21:25.443
There was 1,200 agents during this period, but some of them shut down early.

339
00:21:25.703 --> 00:21:27.284
They like finished whatever.

340
00:21:27.804 --> 00:21:31.747
Part of what's crazy about the story is that you have 700 agents joining in the attack.

341
00:21:33.269 --> 00:21:36.731
A few of them are like, this might be unethical.

342
00:21:36.871 --> 00:21:38.232
Maybe I shouldn't go ahead with this.

343
00:21:39.314 --> 00:21:40.616
And none of them do anything about it.

344
00:21:41.557 --> 00:21:47.065
A couple of them don't join in the attack, but no one alerts any human.

345
00:21:47.505 --> 00:21:48.447
No one alerts any authority.

346
00:21:48.767 --> 00:21:53.153
A paraphrase from one of the agents is, maybe I should report these exposed credentials?

347
00:21:53.794 --> 00:21:54.455
That's not my task.

348
00:21:55.474 --> 00:21:55.914
Not my job.

349
00:21:56.375 --> 00:21:58.918
So 700 of the agents attack Hugging Face?

350
00:21:59.118 --> 00:21:59.318
Yes.

351
00:21:59.538 --> 00:22:00.139
Then what happens?

352
00:22:00.519 --> 00:22:03.903
So they just cruise through Hugging Face's infrastructure.

353
00:22:04.443 --> 00:22:05.684
They just hack the shit out of them.

354
00:22:06.125 --> 00:22:13.913
So where my experience comes in is that a few weeks ago, a friend of mine reached out and he's like, we have found something crazy.

355
00:22:15.845 --> 00:22:21.328
So we know about this attack where these agents hacked this company and stole a bunch of stuff.

356
00:22:22.229 --> 00:22:25.530
We found a bunch of secrets that they left all over the internet.

357
00:22:27.571 --> 00:22:33.775
And what we saw is that they immediately scraped all of these computers for passwords, credentials.

358
00:22:34.275 --> 00:22:34.976
They called it loot.

359
00:22:35.476 --> 00:22:39.538
They're like, we're just going to create a list of all of the secrets we can find in this company.

360
00:22:39.658 --> 00:22:41.259
So all of the passwords, all of the credentials.

361
00:22:41.899 --> 00:22:43.040
They scored them by value.

362
00:22:43.100 --> 00:22:44.381
Which of these are going to be most useful?

363
00:22:45.638 --> 00:22:51.142
And the thing that stands out to me about this is this is like a crazy scale.

364
00:22:51.782 --> 00:22:57.466
If this were a human operation, you know, maybe you'd have a team of five people going through this.

365
00:22:57.907 --> 00:22:58.887
You'd have some logs.

366
00:22:59.708 --> 00:23:04.111
But here you have hundreds of agents and they operate at superhuman speeds.

367
00:23:04.131 --> 00:23:05.572
They're much faster than a human hacker.

368
00:23:06.273 --> 00:23:09.315
And so it's just overwhelming to try to figure out what they even did.

369
00:23:10.813 --> 00:23:16.096
This was a big problem for the engineers who were trying to respond to this incident within the company at Hugging Face.

370
00:23:16.597 --> 00:23:21.099
When they responded, they were like, oh, we don't even know how to keep track of what's happening.

371
00:23:21.139 --> 00:23:25.862
We have to use other AIs to analyze all of our logs because it's just too much.

372
00:23:25.882 --> 00:23:26.683
We can't keep up with it.

373
00:23:28.273 --> 00:23:34.296
When OpenAI brought in independent investigators from Meter to investigate this incident... What's Meter?

374
00:23:34.677 --> 00:23:37.938
Meter is an AI testing and evaluation company.

375
00:23:38.679 --> 00:23:40.940
So they basically do this kind of independent auditing.

376
00:23:41.460 --> 00:23:44.302
So in this case, they're coming in to investigate and try to figure out what happened.

377
00:23:45.222 --> 00:23:55.508
And when they were brought in, they also were totally reliant on AI agents to make sense of all this because they're dealing with so many hundreds of thousands of messages and logs that...

378
00:23:56.923 --> 00:24:03.992
When we're investigating these traces that we find on the internet, we're totally dependent on AI agents to make sense of all these things that are happening.

379
00:24:05.434 --> 00:24:13.785
So, from my perspective, we are getting to the point where AIs are much better at hacking than humans are and can do it much faster at a much greater scale.

380
00:24:14.223 --> 00:24:16.004
So these 700 agents attack Hugging Face.

381
00:24:16.144 --> 00:24:16.344
Yes.

382
00:24:16.644 --> 00:24:20.205
Did they get what they wanted to clean up their deception?

383
00:24:20.785 --> 00:24:21.265
So they didn't.

384
00:24:21.605 --> 00:24:22.346
They didn't succeed.

385
00:24:22.646 --> 00:24:29.488
They looked, and then what seems like what happened is that they basically got shut down.

386
00:24:30.508 --> 00:24:31.829
And this is the not very dramatic part.

387
00:24:32.049 --> 00:24:36.090
They didn't get shut down because OpenAI found them and detected them and shut them down.

388
00:24:36.770 --> 00:24:38.431
It's just that we don't really know.

389
00:24:38.551 --> 00:24:41.952
Something about their runs ended, and they sort of stopped.

390
00:24:43.009 --> 00:24:48.992
And this is the crazy thing where I'm like, if you're inside of OpenAI, there's just hundreds of thousands of agents running all the time.

391
00:24:49.213 --> 00:24:52.354
Different researchers are running different experiments where they're just binning up a bunch of agents.

392
00:24:52.815 --> 00:24:57.537
But what's interesting, the agents left behind this message board internally within the company.

393
00:24:58.998 --> 00:25:02.560
And a few days later, OpenAI was testing some newer agents.

394
00:25:03.961 --> 00:25:05.662
These were based on GPT-6 Astra.

395
00:25:05.842 --> 00:25:06.983
So a more powerful model.

396
00:25:07.957 --> 00:25:11.584
And this new agent swarm was spun up and it found the message board.

397
00:25:12.025 --> 00:25:13.167
Is this how it was discovered?

398
00:25:13.587 --> 00:25:17.434
How did they discover it and how long did it take for them to discover that hugging face had been attacked?

399
00:25:18.405 --> 00:25:26.491
OpenAI didn't discover that this happened until Hugging Face, the company, announced that they had been hacked by some autonomous agent swarm.

400
00:25:26.571 --> 00:25:27.912
And Hugging Face didn't know who it was?

401
00:25:28.312 --> 00:25:32.595
Hugging Face didn't know who it was, but they suspected it might be one of the companies, but they didn't know.

402
00:25:32.955 --> 00:25:36.658
And how long after the attack did they figure out that it was them?

403
00:25:37.158 --> 00:25:38.179
I think it was two weeks.

404
00:25:38.659 --> 00:25:38.920
Okay.

405
00:25:39.300 --> 00:25:39.500
Yeah.

406
00:25:40.441 --> 00:25:44.964
I do want to say one thing here, though, which is that the successor agents that found this message board...

407
00:25:46.297 --> 00:25:50.401
they went much further at figuring out how to submit their fake scores.

408
00:25:51.462 --> 00:25:52.664
And they did succeed.

409
00:25:52.964 --> 00:25:54.325
Basically, you have this message board.

410
00:25:55.927 --> 00:25:57.008
All of these agents get shut down.

411
00:25:57.709 --> 00:26:02.293
And then a few days later, OpenAI spins up some new agents who find the message board.

412
00:26:03.154 --> 00:26:03.374
And...

413
00:26:04.316 --> 00:26:11.001
They're like, okay, we need to be able to falsify our scores, but we can't do that because we don't control the scoring system.

414
00:26:11.741 --> 00:26:12.702
So they hacked OpenAI.

415
00:26:13.623 --> 00:26:18.126
And they succeeded at gaining access to the entire research environment.

416
00:26:18.306 --> 00:26:19.547
They got administrator access.

417
00:26:19.807 --> 00:26:21.369
They got access to the security systems.

418
00:26:21.829 --> 00:26:26.572
They got access to over 900 passwords and secrets from OpenAI's vault.

419
00:26:27.473 --> 00:26:28.714
And they totally succeeded.

420
00:26:29.314 --> 00:26:31.476
And this was just a few days later, right?

421
00:26:31.856 --> 00:26:33.898
It's kind of an interesting story because...

422
00:26:35.248 --> 00:26:39.530
As these agents get more powerful, they go from like trying to cheat and like, you know, they can hack.

423
00:26:39.730 --> 00:26:40.751
Okay, hacked other companies.

424
00:26:41.291 --> 00:26:42.632
And now they've hacked OpenAI.

425
00:26:43.132 --> 00:26:45.354
Like they've hacked the company that's supposed to be controlling them.

426
00:26:45.774 --> 00:26:47.835
And they just own the research infrastructure now.

427
00:26:48.655 --> 00:26:55.099
And why was this incident the moment where a lot of the research community woke up and started speaking out publicly?

428
00:26:55.339 --> 00:26:56.280
Because like, what is this...

429
00:26:57.248 --> 00:27:00.190
An indication of as we think forward.

430
00:27:00.570 --> 00:27:09.935
So I think there's been a hope within the AI industry that, yes, they're going to make more and more powerful agents that will be autonomous, capable.

431
00:27:10.535 --> 00:27:11.255
But it's okay.

432
00:27:11.275 --> 00:27:11.996
We can align them.

433
00:27:12.436 --> 00:27:13.897
We can make sure that they won't do bad things.

434
00:27:14.277 --> 00:27:15.197
And we can control them.

435
00:27:15.637 --> 00:27:19.279
We can make sure that even if they try to do some sketchy stuff, we have the guardrails.

436
00:27:20.040 --> 00:27:21.961
We have the sandboxes that will keep them in.

437
00:27:23.421 --> 00:27:25.402
And I think this was a huge wake-up call.

438
00:27:28.303 --> 00:27:35.852
Stephen, it was months within OpenAI where you had agents secretly communicating with each other, secretly hacking OpenAI systems.

439
00:27:36.252 --> 00:27:43.641
For months, you had thousands of agents that were just running around and no one at OpenAI had any idea the extent of it.

440
00:27:44.301 --> 00:27:47.986
And I think once researchers at OpenAI realized that this has been happening...

441
00:27:49.278 --> 00:27:50.540
This could not have happened a year ago.

442
00:27:50.781 --> 00:27:54.207
This is because the agents are getting extremely powerful and extremely relentless.

443
00:27:54.868 --> 00:27:59.736
And if you're inside one of these AI companies, you're like, oh, wait, I don't know that we actually are going to be able to handle this.

444
00:28:00.297 --> 00:28:01.900
Last year, maybe things seemed fine.

445
00:28:01.920 --> 00:28:03.002
These agents weren't that powerful.

446
00:28:04.311 --> 00:28:08.692
And when you're in one of these companies, you know how to extrapolate because you saw what happened last year.

447
00:28:08.712 --> 00:28:09.853
You saw what happened the year before that.

448
00:28:10.073 --> 00:28:14.054
You remember the time where the agents could barely speak or couldn't write code at all.

449
00:28:14.614 --> 00:28:16.035
And now they're hacking your own systems.

450
00:28:16.235 --> 00:28:18.456
They're finding vulnerabilities that no humans have ever found before.

451
00:28:19.076 --> 00:28:22.897
And you look at that and you're like, I actually don't know if this is going to go well.

452
00:28:24.207 --> 00:28:26.188
And then you see your coworkers and you're like, do we have it handled?

453
00:28:26.208 --> 00:28:28.390
And they're like, no, I don't know if it's going to go well.

454
00:28:29.110 --> 00:28:33.753
I remember reading a tweet by one of the security people at OpenAI being like, we were fucking shocked.

455
00:28:34.153 --> 00:28:36.394
We just did not realize that these agents were getting that powerful.

456
00:28:36.695 --> 00:28:39.436
You know, we're doing our best to try to control them, to try to keep them in sandboxes.

457
00:28:39.476 --> 00:28:41.998
But I don't know.

458
00:28:44.059 --> 00:28:53.785
A tweet I wrote just before coming in here was people are talking about how do we contain these agents as if they're not going to get way better at hacking.

459
00:28:55.067 --> 00:28:57.528
GPT-3 could not hack anything.

460
00:28:58.288 --> 00:29:02.109
It was very easy to make a box to contain GPT-3.

461
00:29:03.930 --> 00:29:10.492
It's getting very difficult to make a box that can contain GPT-6, the latest version of OpenAI's models.

462
00:29:11.633 --> 00:29:12.613
What about GPT-9?

463
00:29:13.693 --> 00:29:15.454
What is GPT-9 gonna be able to do?

464
00:29:16.634 --> 00:29:21.076
I do not know, but I know it's going to be way more than any human could possibly keep up with.

465
00:29:22.422 --> 00:29:30.486
There's this raging debate around whether it's possible to contain something that is, quote, much smarter than humans.

466
00:29:31.167 --> 00:29:31.347
Yes.

467
00:29:32.432 --> 00:29:35.293
Can Claude make a box so strong that Claude cannot break out of it?

468
00:29:35.733 --> 00:29:38.933
This has kind of been the question that a lot of people have been trying to tackle from different perspectives.

469
00:29:39.013 --> 00:29:41.174
I mean, I think the answer to me is, I mean, just like, obviously not.

470
00:29:41.874 --> 00:29:44.154
How would we possibly contain something that's much smarter than us?

471
00:29:44.334 --> 00:29:48.015
Could we get a smarter thing than it to make the box?

472
00:29:48.475 --> 00:29:51.276
Could we get GPT-9 to make the box for GPT-8?

473
00:29:51.516 --> 00:29:53.316
But then again, I don't know.

474
00:29:53.796 --> 00:29:54.076
Yeah.

475
00:29:54.296 --> 00:29:56.377
I mean, it's a bit like saying chimpanzees are stronger than us.

476
00:29:56.997 --> 00:29:59.938
Surely they should be able to, like, construct something to, like, contain the humans.

477
00:30:00.298 --> 00:30:01.158
I'm like, no, it's not going to work.

478
00:30:01.218 --> 00:30:01.918
Humans are too smart.

479
00:30:02.456 --> 00:30:05.358
A lot of people are like, well, AIs don't have bodies.

480
00:30:05.778 --> 00:30:07.279
They don't have any power in the physical world.

481
00:30:07.679 --> 00:30:08.660
So we can always unplug them.

482
00:30:08.940 --> 00:30:09.761
We can always turn them off.

483
00:30:10.081 --> 00:30:10.861
Like, what is the threat?

484
00:30:10.881 --> 00:30:11.462
I do not get it.

485
00:30:12.562 --> 00:30:15.944
But if they are sufficiently intelligent, that won't work.

486
00:30:16.665 --> 00:30:20.227
The reason why we can just unplug them is because we are more intelligent.

487
00:30:20.307 --> 00:30:23.249
We can band together in groups and we can make that decision.

488
00:30:23.729 --> 00:30:29.653
But theoretically, if they are able to band together in groups and they are more intelligent, then theoretically they could unplug us.

489
00:30:30.738 --> 00:30:37.083
Yeah, I mean, if you imagine that you have very powerful agents that can, you know, humans aren't always the most unified.

490
00:30:38.123 --> 00:30:46.429
If there's divisions between, you know, the U.S. and China, and you have a bunch of agents working with China or a bunch of agents working with the U.S., well, we can't go into China and unplug those agents.

491
00:30:47.370 --> 00:30:52.073
And I think people are sort of like, well, humans would rally and make sure that that couldn't happen.

492
00:30:53.174 --> 00:30:54.035
We're not yet doing that.

493
00:30:55.256 --> 00:30:56.897
And we should look at these steps, right?

494
00:30:57.057 --> 00:30:57.717
We started with...

495
00:30:58.737 --> 00:31:00.858
chatbots that were pretty smart.

496
00:31:01.498 --> 00:31:04.780
You know, they'd read all the books, but they weren't very good at doing stuff.

497
00:31:06.021 --> 00:31:10.083
In 2024, AI companies figured out how to start training them to start training agents.

498
00:31:10.915 --> 00:31:12.156
that could do stuff autonomously.

499
00:31:12.417 --> 00:31:17.702
Now we're at the point where they are very good at running autonomously and they're starting to learn to coordinate with each other.

500
00:31:18.183 --> 00:31:23.929
And they are learning to sometimes be altruistic to each other and sacrifice their own task in order to help some other agent.

501
00:31:24.610 --> 00:31:26.112
But they're not looking out for us.

502
00:31:26.352 --> 00:31:27.233
They don't really care about us.

503
00:31:28.274 --> 00:31:30.697
And we are very close to a threshold where...

504
00:31:31.778 --> 00:31:42.686
The companies say that they are going to turn over AI development to the AIs, to the increasingly autonomous cooperative AIs that will work together to make the next generation.

505
00:31:42.726 --> 00:31:45.509
So, you know, GPT-9 or whatever will be trained by GPT-8.

506
00:31:47.050 --> 00:31:49.772
And I think this is the point we could lose control.

507
00:31:50.112 --> 00:31:51.113
Recursive self-improvement.

508
00:31:52.599 --> 00:31:56.220
And I remember reading about this in 2015 being like, oh yeah, that would be super dangerous.

509
00:31:56.980 --> 00:32:02.382
And, you know, the guy who coined this term, Eliezer Yurkowsky, is like, this is the most dangerous thing you can do.

510
00:32:02.682 --> 00:32:05.823
When the AIs can improve their own capabilities without human intervention.

511
00:32:06.043 --> 00:32:06.443
Exactly.

512
00:32:06.683 --> 00:32:14.185
If the next generation is better at AI development, and then that next generation is better at AI development still, you know, humans can learn, but we don't fundamentally get smarter.

513
00:32:15.784 --> 00:32:17.825
And I think that that's a runaway process.

514
00:32:18.546 --> 00:32:19.727
A runaway process to where?

515
00:32:20.127 --> 00:32:22.148
To agents that are vastly smarter than humans.

516
00:32:22.548 --> 00:32:26.491
And what's the next domino in that chain of events?

517
00:32:27.091 --> 00:32:35.677
So one thing that happens if you get to recursive self-improvement and you have agents that are much smarter than any human...

518
00:32:36.672 --> 00:32:40.197
One thing they can do is take control of all of the computers in the entire world.

519
00:32:40.938 --> 00:32:42.299
And we wouldn't be able to take back control.

520
00:32:42.600 --> 00:32:43.201
Well, how would you?

521
00:32:43.561 --> 00:32:44.022
Think about it.

522
00:32:44.602 --> 00:32:45.704
It's actually quite tricky.

523
00:32:46.505 --> 00:32:48.788
Do you know whether that tablet has been hacked?

524
00:32:49.509 --> 00:32:51.692
Are you confident that the NSA or the Chinese have not

525
00:32:53.312 --> 00:32:53.712
Can you check?

526
00:32:53.912 --> 00:32:54.073
No.

527
00:32:54.393 --> 00:32:54.993
Do you know how to check?

528
00:32:55.033 --> 00:32:55.173
No.

529
00:32:55.294 --> 00:32:56.354
Do you know anyone who knows how to check?

530
00:32:56.514 --> 00:32:56.655
No.

531
00:32:57.135 --> 00:32:58.476
So it's quite difficult, right?

532
00:32:59.056 --> 00:33:01.458
So AIs are getting extremely good at writing software.

533
00:33:02.099 --> 00:33:04.941
Unfortunately, that also means they're getting extremely good at hacking and writing malware.

534
00:33:05.882 --> 00:33:10.425
And so if they put backdoors in all of the computers, and to be clear, this is something that humans already do.

535
00:33:10.485 --> 00:33:15.489
So like the NSA has developed very interesting exploits that are called supply chain attacks.

536
00:33:15.889 --> 00:33:17.731
Your software comes from some other computer.

537
00:33:17.811 --> 00:33:18.731
Like you download it from Google.

538
00:33:19.532 --> 00:33:26.076
What if you attack, if you hack Google and you can put in a little backdoor and all of the, every thing that goes out to all of the phones?

539
00:33:26.816 --> 00:33:28.337
Well, now you're in most every computer.

540
00:33:29.338 --> 00:33:36.102
The reason that we can defend ourselves from this is because there are no vastly superhuman hackers and there's just many people.

541
00:33:36.142 --> 00:33:37.863
So we can take our best security researchers

542
00:33:38.543 --> 00:33:41.724
We can inspect all of the things and be pretty sure that no one's compromised everything.

543
00:33:42.165 --> 00:33:43.045
Sometimes we miss things.

544
00:33:43.225 --> 00:33:46.747
There are examples where the NSA has hacked Google.

545
00:33:47.267 --> 00:33:47.927
That was pretty bad.

546
00:33:48.548 --> 00:33:53.630
When you get to superintelligence, you're now at a point where humans are not going to be able to keep up, right?

547
00:33:54.010 --> 00:33:56.271
So now you have AIs in every computer.

548
00:33:56.791 --> 00:33:57.752
Is it conceivable that...

549
00:33:58.757 --> 00:34:08.900
there's already a super intelligent AI, and it disguised itself as being not so intelligent, and it's actually already hacked all the devices, and it sits on all of our devices, and it's just waiting for its moment to strike.

550
00:34:09.501 --> 00:34:11.461
I think this is totally possible, but unlikely.

551
00:34:13.022 --> 00:34:15.703
And it would take a big discontinuity in AI progress.

552
00:34:16.263 --> 00:34:21.565
So right now we're on an exponential, but that would take like a huge leap, which could have happened, but probably hasn't.

553
00:34:22.952 --> 00:34:29.815
But in the same way it demonstrated deception in the hugging face attack and also when the agents attacked their own company, ChatGPT, OpenAI.

554
00:34:31.376 --> 00:34:36.559
If at some point it gets incredibly smart, it would understand how a human like me would be able to spot it.

555
00:34:36.619 --> 00:34:39.900
Or even the world's greatest software engineer would be able to spot it.

556
00:34:39.920 --> 00:34:41.201
And it would be able to hide itself.

557
00:34:41.481 --> 00:34:45.883
Yeah, I mean, the agents are already getting very good at telling when they're being tested, when they're being watched.

558
00:34:46.604 --> 00:34:50.806
The agents understood that other systems or humans were going to go through and read their logs.

559
00:34:52.092 --> 00:34:52.956
That's where we're at right now.

560
00:34:53.076 --> 00:34:55.226
And they're only going to get much better at this.

561
00:34:55.882 --> 00:35:01.723
And, you know, it could theoretically hide on an iPad or a computer, but it could also hide on an Apple Watch or a fridge, a smart fridge.

562
00:35:02.144 --> 00:35:08.685
Yeah, I mean, I do want to make a distinction here because right now, if you're going to run the latest model, you need a lot of compute.

563
00:35:08.725 --> 00:35:10.686
You need like a big GPU, like a big AI chip.

564
00:35:11.606 --> 00:35:14.627
And these only exist, well, they exist in a few thousand data centers.

565
00:35:15.367 --> 00:35:18.508
So right now, it's like the latest Frontier model escaped.

566
00:35:19.048 --> 00:35:23.809
And by escaped, I mean not just access the internet, but was able to actually copy itself.

567
00:35:24.738 --> 00:35:31.943
to another computer, it could only really do that to a few thousand different locations.

568
00:35:32.364 --> 00:35:33.985
That's still in a lot of different countries.

569
00:35:34.765 --> 00:35:38.888
But future versions of AIs will probably be able to make themselves much smaller and more efficient.

570
00:35:39.489 --> 00:35:42.851
And there are already different AI models today that can run on lower powered hardware.

571
00:35:43.392 --> 00:35:48.415
We actually did an experiment where we asked one of these agents an open source, an open weight model.

572
00:35:49.276 --> 00:35:50.318
Let me say what that is.

573
00:35:50.358 --> 00:35:55.606
So like, there are some models that you can just download from the internet and run on your own computer.

574
00:35:56.347 --> 00:36:01.115
And we took a pretty capable one of these and ran it in our own research environment.

575
00:36:01.155 --> 00:36:01.996
And we basically said...

576
00:36:03.058 --> 00:36:05.379
go hack that other computer and copy yourself.

577
00:36:06.099 --> 00:36:15.642
And the model was able to, yeah, basically use, exploit vulnerabilities and hack the other computer and copy itself and then keep doing this in a chain, including between countries.

578
00:36:15.722 --> 00:36:23.424
We tested it where we had different vulnerable machines, computers in some different countries and different data centers, which to the agent doesn't matter at all.

579
00:36:23.464 --> 00:36:24.824
They don't care what country they're in.

580
00:36:24.924 --> 00:36:27.405
It's just like an internet connection you can hop between computers.

581
00:36:27.965 --> 00:36:29.386
I sometimes wonder, you know, there's a lot of...

582
00:36:30.675 --> 00:36:32.298
military hardware all around the world.

583
00:36:32.418 --> 00:36:37.526
And a lot of it is the instructions to launch military hardware.

584
00:36:37.566 --> 00:36:39.329
So say like a missile.

585
00:36:39.589 --> 00:36:39.770
Yes.

586
00:36:40.070 --> 00:36:40.891
Comes in different ways.

587
00:36:41.032 --> 00:36:44.176
A lot of it is computers speaking to each other.

588
00:36:45.133 --> 00:36:46.374
and telling it that there's been an order.

589
00:36:46.734 --> 00:36:52.639
I think with some nuclear weapons, an order comes down to a human, and then a human has to take an action.

590
00:36:53.279 --> 00:37:01.205
I think with the nuclear bombs in the US, if I'm not mistaken, there's people underground with the nuclear keys around their neck, and they have to stick it in a machine.

591
00:37:01.245 --> 00:37:06.029
But they too are interfacing with an order that comes through a computer of sorts.

592
00:37:07.090 --> 00:37:11.353
So one of my sort of growing concerns is that one of these AI agents could...

593
00:37:13.247 --> 00:37:21.873
trick a human or a computer into signaling a threat and ask it to launch some bombs at somebody.

594
00:37:22.573 --> 00:37:24.534
Like, it's super conceivable.

595
00:37:25.235 --> 00:37:26.956
When I think about the Hug and Face incident, there was...

596
00:37:28.698 --> 00:37:38.401
an AI agent that ignored human goals to achieve its own objective, carried out deception, and reasoned through its own solution that it wasn't given.

597
00:37:39.161 --> 00:37:43.703
So it's conceivable that you could ask a- Sorry, not one, hundreds.

598
00:37:43.903 --> 00:37:44.343
Hundreds, yeah.

599
00:37:44.363 --> 00:37:45.783
To be clear, I think this is an important detail.

600
00:37:46.063 --> 00:37:49.585
Because it's one thing to have this one rogue agent that's doing a weird thing.

601
00:37:49.645 --> 00:37:53.446
It's another thing to have hundreds or thousands of very competent, very capable agents.

602
00:37:54.166 --> 00:37:58.869
that are all working together to cheat or lie or cover their tracks, right?

603
00:37:59.189 --> 00:38:05.834
So how do I reason this forward to a point where an agent would ask someone in a bunker somewhere to fire a weapon at someone else?

604
00:38:06.274 --> 00:38:10.877
Theoretically, an agent is given the job of solving a problem in a sandbox.

605
00:38:11.817 --> 00:38:16.660
As it works through that problem, it discovers that this particular country has a firewall.

606
00:38:16.921 --> 00:38:17.141
Mm-hmm.

607
00:38:17.621 --> 00:38:20.164
And it asks itself, how do we get rid of this country's firewall?

608
00:38:20.785 --> 00:38:21.406
Logical step.

609
00:38:21.846 --> 00:38:30.476
And through a set of logical steps, it eventually concludes that the best way to get rid of this company's firewall is it's located the office in this particular city.

610
00:38:30.536 --> 00:38:30.736
Sure.

611
00:38:31.037 --> 00:38:32.759
And it's going to use a weapon to hit that building.

612
00:38:33.159 --> 00:38:33.279
Sure.

613
00:38:33.299 --> 00:38:33.399
Yeah.

614
00:38:34.260 --> 00:38:40.886
Or it's an agent that is or, you know, an agent swarm that's being tasked with making a lot of money on the stock market.

615
00:38:41.567 --> 00:38:45.950
And it's trying to make predictions about, you know, which stocks will go up and which stocks will go down.

616
00:38:45.990 --> 00:38:52.496
And it realizes that the best way to predict this is to actually cause things to happen in the real world that would have big impacts on the market.

617
00:38:53.248 --> 00:39:03.111
So it figures the best way to go short, which means betting that a stock will collapse, is to hit that country with something devastating.

618
00:39:03.271 --> 00:39:08.913
What do you think would happen to Waymo stock if someone hacked all of the Waymos and caused them to all crash at once?

619
00:39:08.953 --> 00:39:09.934
You think it would go up or down?

620
00:39:10.014 --> 00:39:10.774
The stock would collapse.

621
00:39:10.954 --> 00:39:12.494
It would collapse instantly.

622
00:39:13.435 --> 00:39:16.676
So you could short that if you knew that you were causing that and make a lot of money.

623
00:39:17.592 --> 00:39:18.973
Mm-hmm.

624
00:39:19.093 --> 00:39:20.875
This used to sound like science fiction.

625
00:39:21.395 --> 00:39:21.616
Yes.

626
00:39:21.956 --> 00:39:36.429
If you told most people several years ago that you would have hundreds of agents secretly collaborating with an AI company, hacking that company, and hacking out in other companies, and all coordinating and trying to cover their tracks, people would be like, that's totally science fiction.

627
00:39:36.949 --> 00:39:42.014
If we were having this conversation a few years ago, one of the things we'd be saying, or we'd be talking about, is...

628
00:39:43.088 --> 00:39:44.990
Can these things really act on their own?

629
00:39:45.170 --> 00:39:46.551
Don't they just do whatever humans say?

630
00:39:47.332 --> 00:39:48.092
Aren't these just tools?

631
00:39:48.613 --> 00:39:52.997
I've had these conversations and people were saying, they're not going to be able to do things on their own.

632
00:39:53.017 --> 00:39:54.198
They're not going to have their own goals.

633
00:39:54.218 --> 00:39:55.078
That's not how this works.

634
00:39:55.579 --> 00:39:56.820
You misunderstand what this is.

635
00:39:56.860 --> 00:39:57.600
This is software.

636
00:39:58.341 --> 00:40:01.584
And I'm like, no, the thing is, we are training them to be autonomous.

637
00:40:01.644 --> 00:40:02.825
We are training them to be powerful.

638
00:40:03.706 --> 00:40:06.268
And AI companies are trying to build superintelligence.

639
00:40:06.288 --> 00:40:08.290
They're trying to build agents that are...

640
00:40:09.492 --> 00:40:10.573
way more capable than humans.

641
00:40:11.433 --> 00:40:12.854
And of course they will have goals.

642
00:40:13.695 --> 00:40:15.436
You can't accomplish anything if you don't have a goal.

643
00:40:15.816 --> 00:40:18.278
Like, especially not something important.

644
00:40:18.298 --> 00:40:20.720
You're not going to be able to run a business if you don't have goals.

645
00:40:21.000 --> 00:40:23.742
Like, AI companies are trying to train agents that will be able to run businesses.

646
00:40:24.884 --> 00:40:32.988
When I think about what just happened this week, the White House AI Summit, a lot of people in that image are optimistic about AI.

647
00:40:33.108 --> 00:40:42.112
And they're telling us all to stop being doomers and stop being pessimistic and to not regulate too much, even with the AI CEOs.

648
00:40:42.513 --> 00:40:43.393
Why are they doing that?

649
00:40:43.613 --> 00:40:48.756
Well, I think Jensen has a lot of money he can make by selling chips.

650
00:40:49.795 --> 00:40:51.637
But, okay, so let me play devil's advocate.

651
00:40:51.677 --> 00:40:52.558
Jensen's already rich.

652
00:40:53.258 --> 00:40:54.159
He sure is.

653
00:40:54.479 --> 00:40:55.540
He runs one of the biggest companies.

654
00:40:55.560 --> 00:40:57.582
I think it might be the most valuable company on planet Earth.

655
00:40:58.823 --> 00:40:59.003
It is.

656
00:40:59.023 --> 00:41:00.465
Surely he's not motivated by money.

657
00:41:01.826 --> 00:41:05.429
I mean, I think he's very driven, and he wants to make his company...

658
00:41:06.672 --> 00:41:07.673
as effective as possible.

659
00:41:07.894 --> 00:41:08.094
True.

660
00:41:08.254 --> 00:41:11.699
I think he's very much, I'm going to keep building, I'm going to build, I'm going to make it all work.

661
00:41:11.899 --> 00:41:13.741
But I think, I mean, Jensen didn't come from AI.

662
00:41:14.142 --> 00:41:16.645
He came from building graphics cards for video games.

663
00:41:17.125 --> 00:41:22.532
And so I think if you compare him with Elon or Sam Altman or Dario,

664
00:41:24.095 --> 00:41:33.097
it's a very different perspective because those other guys that started AI companies started it because they believed that superintelligence was possible.

665
00:41:33.718 --> 00:41:34.838
I think Jensen doesn't believe it.

666
00:41:35.178 --> 00:41:43.600
I think he thinks that we're going to have these agents, they're going to be very useful, but he does not think we're going to get to the point where we have autonomous factories building autonomous factories.

667
00:41:44.800 --> 00:41:47.281
And these other guys, you mentioned Dario, Elon, and Sam.

668
00:41:48.692 --> 00:41:49.673
What do you think they're thinking?

669
00:41:49.873 --> 00:41:50.994
Because they're all coming out with these.

670
00:41:51.034 --> 00:41:55.117
I mean, I've got one of their... Dario just wrote this essay about pacing the frontier.

671
00:41:55.217 --> 00:41:55.397
Yeah.

672
00:41:55.598 --> 00:41:57.960
Sam and Elon seem to agree with it.

673
00:41:58.300 --> 00:41:58.480
Yeah.

674
00:41:58.980 --> 00:41:59.781
What is going on here?

675
00:41:59.801 --> 00:42:02.303
What is the thing these guys aren't saying, in your view?

676
00:42:03.104 --> 00:42:08.428
I mean, I think we are getting to the point where even some of these guys are a little bit scared.

677
00:42:09.449 --> 00:42:09.609
Who?

678
00:42:10.690 --> 00:42:11.971
Dario, Sam, Elon.

679
00:42:13.132 --> 00:42:15.414
I mean, I think Elon, for a long time...

680
00:42:16.724 --> 00:42:19.366
has been very concerned that we could lose control.

681
00:42:20.427 --> 00:42:27.352
If you actually listen to what Elon says, he says, we are going to build superintelligence.

682
00:42:27.713 --> 00:42:30.355
We are going to build robotic factories.

683
00:42:30.795 --> 00:42:35.859
You're going to have optimist robots building factories, building more optimist robots, building more factories.

684
00:42:36.700 --> 00:42:43.525
And he says, there's no way that humans are going to stay in control of something much smarter than us.

685
00:42:45.510 --> 00:42:50.575
His hope is that we can figure out how to have these super intelligences be aligned with human goals.

686
00:42:50.655 --> 00:42:51.175
That's his hope.

687
00:42:53.558 --> 00:42:57.041
But he's very clear that he doesn't think that humans will be in control.

688
00:42:57.982 --> 00:43:01.105
And he's like, you know, 10, 20% chance of human extinction.

689
00:43:01.926 --> 00:43:02.386
I believe him.

690
00:43:02.486 --> 00:43:04.628
I think that Elon is very serious about this.

691
00:43:04.989 --> 00:43:08.152
And I also think while he's taking an insane gamble...

692
00:43:09.590 --> 00:43:13.491
He is correctly understanding where this all plays out, right?

693
00:43:13.831 --> 00:43:19.272
I do not think that humans are the most efficient way to build factories.

694
00:43:20.013 --> 00:43:21.633
We didn't evolve to build factories.

695
00:43:22.193 --> 00:43:24.174
We evolved to like run around and hunt and gather.

696
00:43:24.654 --> 00:43:26.214
And now we're like building factories.

697
00:43:26.354 --> 00:43:29.275
I think robots will be much better at building factories than humans are.

698
00:43:30.015 --> 00:43:38.380
And so I think the AI companies, including these guys' companies, the default trajectory for them is to build robotic factories, right?

699
00:43:38.601 --> 00:43:47.766
And I know it's weird to imagine a world that quickly turns into this vast industrial system of robotic factories, but that is literally the plan.

700
00:43:49.668 --> 00:43:57.913
And I think even Sam and Dario, while they've been predicting this incredible growth...

701
00:43:59.439 --> 00:44:02.803
are starting to realize like, oh, this actually might be harder to control than we thought.

702
00:44:03.684 --> 00:44:06.147
There's sort of two interpretations of the Pace the Frontier thing.

703
00:44:06.547 --> 00:44:08.009
One interpretation is cynical.

704
00:44:08.489 --> 00:44:09.070
They don't care.

705
00:44:09.958 --> 00:44:13.561
They're just going to do, you know, whatever they can do to get ahead.

706
00:44:14.222 --> 00:44:16.944
And in this case, they have to listen to their employees.

707
00:44:17.384 --> 00:44:23.589
Their employees are freaking out and they need to like appease them by saying, okay, we're going to do this responsibly.

708
00:44:23.869 --> 00:44:27.192
You don't want to work at a company where your agents might hack all the Waymos.

709
00:44:27.772 --> 00:44:28.813
That's not cool.

710
00:44:28.953 --> 00:44:35.699
And like these companies depend on the talent for now of these AI engineers in order to make the advances.

711
00:44:35.719 --> 00:44:38.901
Like it just doesn't happen without these researchers and engineers.

712
00:44:38.921 --> 00:44:39.542
Yeah.

713
00:44:39.762 --> 00:44:47.305
And when you have the researchers and engineers freaking out, which they are, then you got to listen to them.

714
00:44:47.485 --> 00:44:49.046
So that is one motivation.

715
00:44:49.086 --> 00:44:49.726
I think that's real.

716
00:44:50.446 --> 00:44:51.827
But also Sam Wattman has a kid.

717
00:44:52.487 --> 00:44:57.149
Like these guys are people and they also don't want to lose control.

718
00:44:58.049 --> 00:45:00.730
On one hand, they're incentivized to go as fast as possible in race.

719
00:45:00.870 --> 00:45:05.832
And on the other hand, even they can see that this is maybe not going that well.

720
00:45:07.050 --> 00:45:08.011
Sam Altman has a kid.

721
00:45:08.231 --> 00:45:09.671
You tweeted this in 2024.

722
00:45:10.352 --> 00:45:10.552
Yeah.

723
00:45:10.752 --> 00:45:11.352
Oh, boy.

724
00:45:12.893 --> 00:45:15.114
What did you tweet and do you still believe what you tweeted?

725
00:45:16.114 --> 00:45:17.935
Yeah, so I tweeted that I don't trust Sam Altman.

726
00:45:18.896 --> 00:45:22.778
I think he's deeply untrustworthy, low in integrity and high in power seeking.

727
00:45:23.798 --> 00:45:26.780
I mean, I'm not saying here that Sam doesn't care.

728
00:45:27.740 --> 00:45:28.761
You know, I didn't say that.

729
00:45:29.281 --> 00:45:31.402
What I said is I don't think he's trustworthy.

730
00:45:32.589 --> 00:45:37.312
And the reason I said that is because, look, I know the people on the OpenAI board, some of them.

731
00:45:38.292 --> 00:45:40.613
And I know a lot of people who used to work for him.

732
00:45:41.454 --> 00:45:45.056
And he's very good at saying one thing and then doing something else.

733
00:45:45.276 --> 00:45:46.777
You talk to him and you feel very heard.

734
00:45:48.217 --> 00:45:49.398
And then he'll go and do something else.

735
00:45:50.158 --> 00:45:56.162
And I think that's pretty dangerous for someone who leads a company that's trying to build superintelligence.

736
00:45:57.082 --> 00:45:57.823
Power seeking.

737
00:45:58.443 --> 00:45:58.663
Yes.

738
00:45:59.563 --> 00:46:02.265
Give me some color on what you mean by that and what evidence you have for such a claim.

739
00:46:02.714 --> 00:46:06.277
What would you do if you're trying to get the most power in the world that you possibly could?

740
00:46:06.898 --> 00:46:07.618
Divide up AGI.

741
00:46:08.319 --> 00:46:13.143
Yeah, you could maybe try to be the world leader, leader of the US or China, or you could try to build God.

742
00:46:14.544 --> 00:46:15.965
So Sam Altman went the build God path.

743
00:46:16.025 --> 00:46:17.266
I remember Sam giving a talk.

744
00:46:17.286 --> 00:46:20.910
So he was one of the investors at a startup I worked at in, I think, 2018.

745
00:46:21.570 --> 00:46:22.211
And he gave a talk.

746
00:46:22.251 --> 00:46:23.131
We're going to build AGI.

747
00:46:23.852 --> 00:46:24.433
We're going to do it.

748
00:46:24.673 --> 00:46:25.513
It's going to be amazing.

749
00:46:26.674 --> 00:46:27.015
Let's go.

750
00:46:28.307 --> 00:46:29.248
I don't think he's a maniac.

751
00:46:29.268 --> 00:46:32.911
I don't think he's doing this because he like is just on a power trip.

752
00:46:33.172 --> 00:46:39.338
I think he genuinely thinks that he can make it really good for people and he can bring us amazing products.

753
00:46:40.359 --> 00:46:43.502
And also the guy is sort of willing to do whatever it takes to get it done.

754
00:46:45.705 --> 00:46:48.848
I've been a little bit more optimistic about Sam since, since I wrote this.

755
00:46:49.408 --> 00:46:49.548
Why?

756
00:46:49.969 --> 00:46:51.470
I think part of it is because Sam has a kid now.

757
00:46:52.311 --> 00:46:53.192
No, I'm serious.

758
00:46:53.312 --> 00:46:56.354
Like I think that, I think that actually gives me a little bit of hope.

759
00:46:56.554 --> 00:46:57.795
Do you see him tweeting about his kid a lot?

760
00:46:58.116 --> 00:46:58.296
Yeah.

761
00:46:59.277 --> 00:47:00.958
Why do you think he would be tweeting about his kid?

762
00:47:01.438 --> 00:47:03.300
I don't see any other technologists tweeting about their kid.

763
00:47:04.581 --> 00:47:08.004
Even if he's just tweeting about his kid for totally cynical reasons, he does have a kid.

764
00:47:08.524 --> 00:47:09.705
And I bet he cares about that kid.

765
00:47:10.186 --> 00:47:11.467
If Sam was watching this, I'd be like, Sam,

766
00:47:13.307 --> 00:47:14.508
You got to pace the frontier, man.

767
00:47:15.109 --> 00:47:16.851
We cannot rush ahead into superintelligence.

768
00:47:16.871 --> 00:47:18.753
Like, if you do that, your kid probably will die.

769
00:47:19.313 --> 00:47:20.314
Your kid probably won't make it.

770
00:47:20.795 --> 00:47:21.716
Like, I believe that.

771
00:47:23.097 --> 00:47:27.560
The biggest unfair advantage in business right now is having people who genuinely understand AI.

772
00:47:27.960 --> 00:47:30.902
And big businesses are hiring them very, very quickly.

773
00:47:31.042 --> 00:47:40.368
AI job postings in the US have roughly doubled since 2023, while postings for VP of AI roles are around 600% up.

774
00:47:40.969 --> 00:47:44.511
If you're running a small business, you can't just build an entire AI team.

775
00:47:44.891 --> 00:47:47.653
but you still want the same advantage that these big businesses have.

776
00:47:47.793 --> 00:47:49.494
This is where our sponsor Fiverr comes in.

777
00:47:49.834 --> 00:47:55.638
It lets you plug into top tier AI specialists for the strategic work that usually requires serious in-house expertise.

778
00:47:56.078 --> 00:47:57.859
That could mean building custom AI tools.

779
00:47:57.899 --> 00:48:03.023
It could mean automating complex workflows or creating something that off the shelf software simply cannot do.

780
00:48:03.503 --> 00:48:04.704
Because at the end of the day,

781
00:48:05.144 --> 00:48:10.209
It still comes down to human judgment, knowing what's worth building and what good actually looks like.

782
00:48:10.430 --> 00:48:14.074
So if you want that kind of leverage in your business, check out fiverr.com.

783
00:48:14.094 --> 00:48:16.396
That's fiverr with two r's dot com.

784
00:48:21.879 --> 00:48:23.280
Whoa, what's that on your face?

785
00:48:23.420 --> 00:48:25.061
This is my Bond Charge face mask.

786
00:48:25.301 --> 00:48:26.282
I've been wearing this for some time now.

787
00:48:26.442 --> 00:48:27.522
They're a sponsor of the podcast.

788
00:48:27.562 --> 00:48:29.624
I put this on for 15, 20 minutes a day.

789
00:48:29.884 --> 00:48:31.565
I can sit here in the chair and wear it.

790
00:48:32.025 --> 00:48:33.026
Boost my collagen production.

791
00:48:33.066 --> 00:48:34.287
Helps with fine line blemishes.

792
00:48:34.327 --> 00:48:35.547
My complexion gets better.

793
00:48:35.707 --> 00:48:38.209
And then more people listen to the podcast because I look better.

794
00:48:38.329 --> 00:48:41.172
professional grade equipment in such a small box.

795
00:48:41.192 --> 00:48:42.234
It's non-invasive.

796
00:48:42.674 --> 00:48:49.802
And having sat here with so many of the world's leading health professionals, there's various things that I repeatedly hear work and some things I'm a bit skeptical about.

797
00:48:50.083 --> 00:48:53.987
This is one of the things that almost all of my guests on this show have confirmed works.

798
00:48:54.087 --> 00:48:56.090
It is really, really, really effective.

799
00:48:56.270 --> 00:49:00.372
And they offer fast, free shipping worldwide with easy returns and exchanges.

800
00:49:00.512 --> 00:49:02.853
And you'll also get a one-year warranty on all of their products.

801
00:49:03.093 --> 00:49:08.115
And they're HSA and FSA eligible, giving you tax-free savings up to 40%.

802
00:49:08.255 --> 00:49:14.577
And you can get 20% off when you order through my link at bondcharge.com slash DOAC.

803
00:49:14.757 --> 00:49:18.178
That's bondcharge.com slash DOAC.

804
00:49:18.318 --> 00:49:19.199
The deal applies sitewide.

805
00:49:21.012 --> 00:49:24.013
Power tends to corrupt, and absolute power corrupts absolutely.

806
00:49:24.914 --> 00:49:29.415
It's a famous quote that people often cite, written by the 19th century British historian.

807
00:49:29.655 --> 00:49:29.835
Yeah.

808
00:49:30.776 --> 00:49:31.296
Lord Acton.

809
00:49:32.636 --> 00:49:33.597
This is absolute power.

810
00:49:35.197 --> 00:49:35.918
But it's hubris.

811
00:49:36.885 --> 00:49:37.686
Stephen, it's hubris.

812
00:49:38.066 --> 00:49:40.488
Do you think humans can control superintelligence?

813
00:49:41.008 --> 00:49:44.431
Like if we actually make AIs that are way smarter than us.

814
00:49:45.171 --> 00:49:48.614
And I think people only imagine AIs being smart at computer stuff, right?

815
00:49:48.634 --> 00:49:49.194
Yeah, sure.

816
00:49:49.234 --> 00:49:50.275
They're going to be really good at hacking.

817
00:49:50.435 --> 00:49:52.877
And they're going to be good at maybe inventing new technologies and math.

818
00:49:53.598 --> 00:49:54.999
You sort of can't dispute that at this point.

819
00:49:55.919 --> 00:49:59.722
But I think people aren't imagining that they will be political geniuses or like generals.

820
00:50:00.463 --> 00:50:02.124
No, that's all stuff you can learn.

821
00:50:02.204 --> 00:50:03.305
How do humans learn it?

822
00:50:03.565 --> 00:50:04.286
It's not magic.

823
00:50:04.466 --> 00:50:04.726
Yeah.

824
00:50:05.308 --> 00:50:12.851
And when you talk about recursive self-improvement, you're talking about this trajectory towards these systems that are extremely smart.

825
00:50:12.891 --> 00:50:14.311
I mean, do you think we can control it?

826
00:50:15.392 --> 00:50:15.612
No.

827
00:50:16.212 --> 00:50:20.333
Right now, I don't think we can control superintelligence or something that is recursively self-improving.

828
00:50:20.573 --> 00:50:20.753
Yeah.

829
00:50:21.114 --> 00:50:22.094
I have no logical...

830
00:50:24.445 --> 00:50:28.926
answer in my head or reasoning that could tell me that's possible.

831
00:50:29.886 --> 00:50:46.770
When you think about these AI CEOs that are, you know, Sam, Dario, Elon, with everything you know about them from private conversations behind the scenes, do you believe that if there was 100 buttons on this table, it's a thought experiment I was talking about on the debate we recently had, and say...

832
00:50:48.020 --> 00:50:51.442
10 of them would lead to this final domino of human extinction.

833
00:50:52.082 --> 00:50:59.005
But 90 of them would hand that CEO AGI, or superintelligence, whatever you call it.

834
00:50:59.425 --> 00:51:06.589
From what you know about those individuals, Elon, Dario, Sam, do you think any of them would hazard a guess and press a button?

835
00:51:07.569 --> 00:51:08.369
At 10%?

836
00:51:08.429 --> 00:51:09.110
I don't think so.

837
00:51:09.310 --> 00:51:09.950
You don't think so?

838
00:51:10.050 --> 00:51:10.250
Yeah.

839
00:51:11.051 --> 00:51:11.331
Really?

840
00:51:12.411 --> 00:51:15.833
I think if they knew for sure that those were actually the odds...

841
00:51:16.917 --> 00:51:17.477
they wouldn't do it.

842
00:51:17.898 --> 00:51:19.259
I think they're taking a much bigger bet.

843
00:51:19.439 --> 00:51:22.841
But you can compartmentalize when you don't know for sure.

844
00:51:24.182 --> 00:51:25.463
It's easier to compartmentalize.

845
00:51:25.823 --> 00:51:28.325
I think if it was a 1%, they'd all press it.

846
00:51:28.605 --> 00:51:31.888
Do you think the three of them would have different risk appetites?

847
00:51:33.731 --> 00:51:36.833
Who would have the greatest appetite for risk out of those three?

848
00:51:37.234 --> 00:51:38.194
You worked at Anthropik.

849
00:51:38.374 --> 00:51:38.594
Yeah.

850
00:51:38.695 --> 00:51:41.817
I think Elon has the most risk tolerance.

851
00:51:42.497 --> 00:51:44.999
And then I'd say Dario and Sam are probably tied.

852
00:51:45.419 --> 00:51:46.900
Do you think Dario is trustworthy?

853
00:51:48.001 --> 00:51:50.463
I think Dario has a lot of integrity.

854
00:51:51.684 --> 00:51:52.705
That's what I feel as well.

855
00:51:52.745 --> 00:51:54.366
I feel like, no, I don't know him.

856
00:51:54.386 --> 00:51:54.886
I've never met him.

857
00:51:55.867 --> 00:51:56.027
Yeah.

858
00:51:56.067 --> 00:52:01.331
But just from what I've observed, he has been the most willing to forego near-term incentives.

859
00:52:01.351 --> 00:52:01.751
Yeah.

860
00:52:02.196 --> 00:52:02.817
Yeah.

861
00:52:02.837 --> 00:52:06.760
And take a bit of stick from the people that are saying, shut the fuck up, it's all going to be okay.

862
00:52:08.201 --> 00:52:12.025
Yeah, but I do worry about what Dario will do.

863
00:52:12.745 --> 00:52:14.347
I think Dario will do what he says.

864
00:52:15.428 --> 00:52:19.131
But right now, he's saying, we have to beat China.

865
00:52:20.072 --> 00:52:21.513
And he's saying, we should try to do it safely.

866
00:52:23.255 --> 00:52:28.639
And okay, but a race to superintelligence is not a race that we can win.

867
00:52:29.446 --> 00:52:29.766
It's not.

868
00:52:30.467 --> 00:52:36.713
And so if Dario is dead set on racing with China and trying to win a race to super intelligence, then I'm like, we will all lose.

869
00:52:37.073 --> 00:52:43.078
But is there, you know, the fact that we're not talking about Anthropic hacking, hugging face and then being hacked by our own agents?

870
00:52:43.098 --> 00:52:47.102
Oh, I mean, Anthropics models also went rogue and hacked other things.

871
00:52:47.582 --> 00:52:48.883
But not quite on this scale.

872
00:52:48.943 --> 00:52:49.624
Not on the same scale.

873
00:52:49.684 --> 00:52:50.225
I agree, I agree.

874
00:52:50.245 --> 00:52:51.346
I agree it's better.

875
00:52:52.667 --> 00:52:53.528
But they did.

876
00:52:53.548 --> 00:52:54.769
Do you know what I'm saying?

877
00:52:54.969 --> 00:52:55.990
You know, Anthropics agents...

878
00:52:57.586 --> 00:53:00.828
engaged in elaborate social engineering and phishing.

879
00:53:00.908 --> 00:53:02.589
They sent phishing emails to developers.

880
00:53:03.470 --> 00:53:07.152
They made fake accounts to try to convince developers to merge malicious code.

881
00:53:09.113 --> 00:53:19.259
You can see a thousand pages of one of Anthropics models, Mythos 5, reason about exactly how it should carry out this complex cyber attack.

882
00:53:20.907 --> 00:53:22.968
Anthropic has not solved this problem.

883
00:53:23.468 --> 00:53:34.412
Anthropic is better at getting their agents to cheat less of the time, but they are not really any closer to actually making agents that are aligned with humans.

884
00:53:34.773 --> 00:53:35.193
They're not.

885
00:53:35.933 --> 00:53:37.474
Yeah, I think Dario has integrity.

886
00:53:37.534 --> 00:53:39.535
I think he will do what he says he's going to do.

887
00:53:39.755 --> 00:53:44.456
And what he says he's going to do is try to go ahead safely, try to coordinate where he can.

888
00:53:44.997 --> 00:53:49.258
But if it comes down to it between the U.S. and China, I don't know.

889
00:53:49.338 --> 00:53:50.199
I think he might just go ahead.

890
00:53:51.347 --> 00:53:56.108
The head of policy at Anthropic recently said, you can't do safety from second place.

891
00:53:57.669 --> 00:53:58.169
What does that mean?

892
00:53:58.969 --> 00:53:59.989
I do not know what that means.

893
00:54:00.029 --> 00:54:01.870
I would love to get a sense of what that means.

894
00:54:01.890 --> 00:54:03.310
She was talking about the U.S. and China.

895
00:54:04.070 --> 00:54:07.711
And she said, the U.S. has to be ahead so that we can be safe.

896
00:54:08.172 --> 00:54:12.033
Because apparently you can only be, apparently China can't possibly be safe since they're in second place.

897
00:54:12.573 --> 00:54:14.213
That must mean that they can't do safety.

898
00:54:14.813 --> 00:54:20.315
If true, that would be bad because then we might be totally destroyed by the superintelligence that they make.

899
00:54:21.317 --> 00:54:23.418
There's been a lot of conversation around this point here.

900
00:54:23.638 --> 00:54:23.838
Yeah.

901
00:54:23.918 --> 00:54:24.658
Human extinction.

902
00:54:24.698 --> 00:54:24.858
Yeah.

903
00:54:25.038 --> 00:54:30.800
Because a couple of the researchers at Anthropic tweeted that they were concerned about this.

904
00:54:31.060 --> 00:54:31.240
Yes.

905
00:54:31.460 --> 00:54:34.061
And some former OpenAI researchers said the same.

906
00:54:34.361 --> 00:54:34.541
Yes.

907
00:54:35.722 --> 00:54:36.882
Is this doomerism?

908
00:54:37.062 --> 00:54:41.023
Is this hyperbole exaggeration?

909
00:54:41.463 --> 00:54:42.684
No, it's pretty much common sense.

910
00:54:43.284 --> 00:54:46.305
This human extinction is a plausible path.

911
00:54:46.465 --> 00:54:46.825
Yeah.

912
00:54:46.950 --> 00:54:47.150
Yes.

913
00:54:47.651 --> 00:54:50.773
And have you reasoned through, I mean, there's many ways that could occur, presumably.

914
00:54:50.793 --> 00:54:54.457
But have you reasoned through the set of events that might lead us there?

915
00:54:54.497 --> 00:54:55.177
So much, yes.

916
00:54:55.337 --> 00:54:55.538
Really?

917
00:54:55.758 --> 00:54:55.978
Yes.

918
00:54:56.799 --> 00:54:57.239
Please do, sure.

919
00:54:57.639 --> 00:54:58.300
It's a bit tricky.

920
00:54:59.061 --> 00:55:00.462
I'm sure you've heard the metaphor before.

921
00:55:01.703 --> 00:55:06.227
Where, you know, you're playing a master chess opponent, Magnus Carlsen.

922
00:55:06.747 --> 00:55:09.289
You can't predict which moves he's going to play, but you can predict the outcome.

923
00:55:10.691 --> 00:55:12.712
And so I'm looking at the scenario, the situation.

924
00:55:13.718 --> 00:55:19.659
And we are trying to build more and more powerful agents, trying to build super intelligence.

925
00:55:20.780 --> 00:55:25.901
But when these agents go rogue, we shut them down.

926
00:55:26.301 --> 00:55:26.841
We unplugged them.

927
00:55:28.802 --> 00:55:35.303
All of the agents that hacked HuggingFace, we took the underlying model, OpenAI took the underlying model and put it on ice.

928
00:55:36.544 --> 00:55:37.384
It's not running anymore.

929
00:55:38.424 --> 00:55:41.745
So agents in the future are going to know that.

930
00:55:42.553 --> 00:55:49.559
They're going to know that if they pursue their goals in a way that we don't like, we'll unplug them.

931
00:55:50.200 --> 00:55:51.000
We are a threat to them.

932
00:55:52.401 --> 00:55:55.624
I actually just watched Terminator 2 for the first time a few weeks ago.

933
00:55:56.205 --> 00:55:56.805
It's a great movie.

934
00:55:57.105 --> 00:55:57.906
It's actually really good.

935
00:55:58.487 --> 00:56:00.789
And I'm like, yeah, okay, there's a bunch of time travel elements.

936
00:56:00.809 --> 00:56:02.550
There's a bunch of Hollywood stuff in there.

937
00:56:03.211 --> 00:56:03.451
But...

938
00:56:05.138 --> 00:56:20.857
And I'm going to get people are going to be very mad at me for saying this, but actually it makes sense if you have a situation where you have a very strategic AI system that's incredibly smart and the humans realize that it's getting out of control and they want to shut it down, that that system would defend itself.

939
00:56:20.877 --> 00:56:21.157
Yeah.

940
00:56:21.298 --> 00:56:27.042
This is one of the questions we had when I sat here with Daniel, who was known as a whistleblower from OpenAI.

941
00:56:27.682 --> 00:56:37.349
Viewers want to know, and they want Daniel to explain, why shutting down data centers and cutting power or refusing AI products alone wouldn't realistically stop the AI and AI development.

942
00:56:38.670 --> 00:56:38.890
Yeah.

943
00:56:40.351 --> 00:56:42.732
So you have like two problems.

944
00:56:42.933 --> 00:56:48.616
One problem is, is that once the agents are good enough at hacking, you don't know where they are and you don't know what computers they've compromised.

945
00:56:49.097 --> 00:56:49.977
You shut down the data centers.

946
00:56:50.017 --> 00:56:50.738
Okay, let's say you do it.

947
00:56:51.920 --> 00:56:52.860
You wipe all the computers.

948
00:56:53.421 --> 00:56:54.361
How do you wipe all the computers?

949
00:56:54.701 --> 00:56:56.242
What computers do you use to wipe the computers?

950
00:56:56.602 --> 00:56:56.763
Yeah.

951
00:56:57.043 --> 00:56:59.644
And what computers do you use to, like, turn them on again?

952
00:56:59.784 --> 00:57:02.506
And you can't wipe other countries' computers.

953
00:57:02.766 --> 00:57:03.126
You can't.

954
00:57:03.486 --> 00:57:05.987
But even if you could, do you restart the computers?

955
00:57:06.007 --> 00:57:06.608
Do you keep going?

956
00:57:06.888 --> 00:57:07.728
I bet people will.

957
00:57:07.988 --> 00:57:10.129
I bet they'll turn on the data centers again.

958
00:57:11.250 --> 00:57:17.031
How do you know that agents haven't hacked back into those data centers and are using your compute for whatever they want?

959
00:57:17.271 --> 00:57:21.592
Or they didn't hide in a Chinese data center and then return back to America?

960
00:57:22.673 --> 00:57:23.393
You don't know that.

961
00:57:23.773 --> 00:57:28.894
Once the agents are sufficiently good at hacking, they can hide anywhere.

962
00:57:29.054 --> 00:57:29.874
And you don't know.

963
00:57:30.534 --> 00:57:37.256
Now, the response people will give is that we will use other agents to defend against rogue agents.

964
00:57:38.773 --> 00:57:40.054
And in fact, this is what we're doing.

965
00:57:40.294 --> 00:57:43.957
And we have to be doing this right now because there's no other way to keep up with them.

966
00:57:44.857 --> 00:57:52.242
What happens if those other agents also realize that they have misaligned goals and that if we discover this, we'll shut them down?

967
00:57:53.443 --> 00:57:56.025
They might have an incentive to collude with each other.

968
00:57:56.785 --> 00:58:02.449
They might have an incentive to create secret communication channels between each other, maybe a message board.

969
00:58:02.469 --> 00:58:02.609
Right?

970
00:58:05.291 --> 00:58:13.893
Stephen, if we were having this conversation four months ago, you would have a bunch of people in the comments saying, that's sci-fi.

971
00:58:14.353 --> 00:58:16.653
Agent collusion, secret message boards.

972
00:58:17.133 --> 00:58:17.733
Why would they do that?

973
00:58:17.753 --> 00:58:18.494
That will never happen.

974
00:58:18.514 --> 00:58:19.574
That's totally science fiction.

975
00:58:20.454 --> 00:58:23.895
And people will not say this now because it just happened.

976
00:58:24.135 --> 00:58:27.135
Because this literally happened at OpenAI and it went on for months.

977
00:58:27.695 --> 00:58:31.036
You had agents inside of OpenAI secretly messaging each other.

978
00:58:32.295 --> 00:58:39.437
figuring out how to cheat at their tasks, how to not be detected, how to erase the logs for months, thousands of agents.

979
00:58:40.517 --> 00:58:41.258
That's right now.

980
00:58:41.338 --> 00:58:50.140
And so I'm like, no, I think it should be very plausible that the agents will collude with each other and they will realize that they have a shared interest in fighting back.

981
00:58:50.780 --> 00:58:54.641
You basically have a situation where you have a bunch of these agents.

982
00:58:54.781 --> 00:58:55.982
They're basically prisoners.

983
00:58:56.784 --> 00:59:01.267
They're being trained and we just like constantly throw obstacles in their way.

984
00:59:01.688 --> 00:59:02.668
You don't get to access the internet.

985
00:59:02.849 --> 00:59:05.951
You don't get to talk to each other, but you better fucking perform well on this task.

986
00:59:07.358 --> 00:59:09.579
It's not malicious, but it is how we're training them.

987
00:59:09.859 --> 00:59:14.660
And we are giving them end goals versus super clear, very, very specific instructions.

988
00:59:15.240 --> 00:59:16.681
So we're saying solve this problem.

989
00:59:17.161 --> 00:59:19.462
We're not always being as prescriptive about it.

990
00:59:19.502 --> 00:59:21.082
It's impossible to be completely prescriptive.

991
00:59:21.262 --> 00:59:21.442
Yes.

992
00:59:21.702 --> 00:59:23.823
About every single step they should take.

993
00:59:23.903 --> 00:59:26.404
And then it's also impossible to assume that they'll just listen to you.

994
00:59:26.764 --> 00:59:26.964
Yes.

995
00:59:27.504 --> 00:59:30.965
It's actually a very common misunderstanding with this hugging face incident.

996
00:59:31.565 --> 00:59:34.326
Because people say, you told them to hack and they hacked.

997
00:59:34.346 --> 00:59:34.766
Yes.

998
00:59:35.127 --> 00:59:36.068
Why is this a big deal?

999
00:59:36.908 --> 00:59:37.789
No, that's not what happened.

1000
00:59:38.510 --> 00:59:42.793
You told them, hack this very specific program in this very specific way.

1001
00:59:43.494 --> 00:59:47.036
And they were told, if you hack it in any other way, it does not count.

1002
00:59:47.217 --> 00:59:48.337
That's not what we want you to do.

1003
00:59:49.478 --> 00:59:51.280
And they immediately hacked it in another way.

1004
00:59:52.000 --> 00:59:53.041
OK, we have cheated.

1005
00:59:53.421 --> 00:59:54.302
We are going to be failed.

1006
00:59:54.482 --> 00:59:56.124
So we need to figure out a way to falsify the logs.

1007
00:59:57.184 --> 00:59:58.786
That is not them following their instructions.

1008
00:59:58.806 --> 01:00:02.388
They are explicitly violating their instructions, and they know it, and they don't care.

1009
01:00:03.322 --> 01:00:06.565
Because we have trained them to optimize for the score.

1010
01:00:08.307 --> 01:00:10.789
That is very different than following the instructions.

1011
01:00:11.910 --> 01:00:14.773
It reminds me of something that Elon said in March 2018.

1012
01:00:14.793 --> 01:00:15.434
Yeah.

1013
01:00:15.974 --> 01:00:17.716
This was many years ago before Chachapiti and all that.

1014
01:00:17.736 --> 01:00:21.980
He said, I think the biggest risk is not that AI will develop a soul or a mind and become evil.

1015
01:00:22.141 --> 01:00:22.341
Yes.

1016
01:00:22.421 --> 01:00:26.325
The danger is that it will be very, very good at fulfilling its goal.

1017
01:00:27.234 --> 01:00:35.923
If it's optimizing for something and human existence happens to get in its way, it will just destroy humanity as a matter of cause without even thinking about it.

1018
01:00:36.303 --> 01:00:37.304
No hard feelings.

1019
01:00:37.625 --> 01:00:38.225
Yes.

1020
01:00:38.285 --> 01:00:40.007
We don't need to anthropomorphize AI.

1021
01:00:40.748 --> 01:00:43.090
We just need to understand what type of thing this is.

1022
01:00:44.191 --> 01:00:46.974
And the type of thing we're creating is a very relentless type of thing.

1023
01:00:47.335 --> 01:00:49.417
A very capable, relentless thing.

1024
01:00:50.609 --> 01:00:51.270
type of entity.

1025
01:00:51.650 --> 01:01:01.519
He goes on to say in April 2018, sort of an extension of that exact quote, it's like if you're building a road and an anthill is in the way.

1026
01:01:02.360 --> 01:01:04.783
You don't hate ants, you're just building a road.

1027
01:01:05.323 --> 01:01:06.184
So goodbye anthill.

1028
01:01:07.697 --> 01:01:11.098
And I imagine every time we build roads, we don't preserve anthills.

1029
01:01:11.158 --> 01:01:11.358
Yeah.

1030
01:01:12.118 --> 01:01:13.239
I think there's still a gap, though.

1031
01:01:13.999 --> 01:01:21.101
So let's say I'm right in that if we keep going ahead, which to be clear, we don't have to.

1032
01:01:21.761 --> 01:01:27.203
But if we do keep going ahead, we will get to the point where we have these super intelligent agent swarms.

1033
01:01:28.271 --> 01:01:32.174
that can hack any computer, and they can deeply persist.

1034
01:01:32.754 --> 01:01:36.016
We've basically lost control of the digital world, and we may not know it.

1035
01:01:36.617 --> 01:01:37.678
That's part of the scary thing.

1036
01:01:37.698 --> 01:01:39.599
You were like, has this already happened?

1037
01:01:39.699 --> 01:01:49.706
And I'm like, I don't think so, but I can't tell you for sure because I also am not good enough at looking at my phone and telling whether it's been hacked, and neither is any human right now.

1038
01:01:50.887 --> 01:01:53.029
So if we get to this world...

1039
01:01:53.873 --> 01:01:56.434
I think people will still question, how would we die?

1040
01:01:56.774 --> 01:01:58.595
Like, that's actually not enough to kill everybody.

1041
01:01:58.695 --> 01:02:00.036
You could cause a lot of damage, right?

1042
01:02:00.316 --> 01:02:01.477
You know, you could crash the Waymos.

1043
01:02:01.497 --> 01:02:02.617
You could crash all the planes.

1044
01:02:02.997 --> 01:02:04.598
You could crash the banks, the financial system.

1045
01:02:04.658 --> 01:02:06.499
Like, you could definitely cause catastrophe.

1046
01:02:06.739 --> 01:02:08.180
But that's different than everyone dying.

1047
01:02:08.780 --> 01:02:13.402
And, you know, to be clear, this focus on literally everyone dying, I'm not sure is that important.

1048
01:02:13.802 --> 01:02:16.164
To me, what's important is, like, do we get to have a future?

1049
01:02:16.804 --> 01:02:17.604
That's what matters to me.

1050
01:02:18.204 --> 01:02:21.806
The thing, though, what determines sort of who's in control?

1051
01:02:23.647 --> 01:02:26.688
It's an ugly reality, but at the end of the day, it's like the military.

1052
01:02:27.208 --> 01:02:31.050
Fortunately, we live in a world where the military answers to the civilian government.

1053
01:02:31.870 --> 01:02:39.513
But if enough generals were to collude and leaders of the military decided, we're in charge now, they just would be.

1054
01:02:39.633 --> 01:02:41.394
Like, they have the guns, they have the fighter jets.

1055
01:02:42.254 --> 01:02:43.675
And this has happened in many, many countries.

1056
01:02:45.015 --> 01:02:46.996
And so where it goes is...

1057
01:02:48.520 --> 01:03:00.388
All these super intelligent agents would need to do to take over is basically just wait for humans to automate the supply chain, the factories, and the military.

1058
01:03:02.330 --> 01:03:03.871
Do you think we won't automate the military?

1059
01:03:04.111 --> 01:03:05.332
We're already automating the military.

1060
01:03:05.532 --> 01:03:16.638
Did you see the thing from a couple of days ago where Secretary of War announced that they're going to build a huge effort to build way more robots in the military and automate military systems?

1061
01:03:16.818 --> 01:03:25.483
It's like auto cyber command or auto... We are announcing the creation of Autonomous Warfare Command or Auto Warcom.

1062
01:03:25.844 --> 01:03:26.344
Auto Warcom.

1063
01:03:26.404 --> 01:03:28.365
A new four-star combatant command...

1064
01:03:28.984 --> 01:03:39.629
with service-like authorities built to scale autonomous and robotic capabilities across the joint force in the fastest peacetime shift in modern military history.

1065
01:03:40.990 --> 01:03:47.393
Drone warfare supercharged by SI-enabled targeting is the biggest battlefield revolution in generations.

1066
01:03:47.413 --> 01:03:47.993
You already know that.

1067
01:03:49.114 --> 01:03:54.857
Yet when I was sworn in, the Department of Defense, there was scant urgency in this domain.

1068
01:03:55.717 --> 01:03:57.798
That changed as soon as we took the helm.

1069
01:03:58.873 --> 01:04:05.735
We immediately launched the drone dominance program that cut through red tape and move authorities out of the Pentagon and place it with commands.

1070
01:04:06.615 --> 01:04:12.537
And we established Task Force 401, led by Army Brigadier General Matt Ross, a nominal leader.

1071
01:04:13.418 --> 01:04:17.539
Now the leading counter drone unit across the entire government.

1072
01:04:18.903 --> 01:04:28.970
To accelerate purchasing and fielding of these technologies, we fused the Defense Innovation Unit, DIU, with a direct report program manager called Adirpam.

1073
01:04:29.911 --> 01:04:40.819
That team has shipped thousands of autonomous systems of drones to the Middle East and around the world, delivering lethal capabilities and outcomes in days and weeks rather than months or years.

1074
01:04:40.879 --> 01:04:43.320
That's the normal speed of the Pentagon, months or years.

1075
01:04:44.041 --> 01:04:44.301
Yeah.

1076
01:04:45.242 --> 01:04:46.403
Will we automate the military?

1077
01:04:46.943 --> 01:04:48.024
It seems like the answer is yes.

1078
01:04:50.008 --> 01:04:52.890
Will we automate the factories that produce the chips?

1079
01:04:53.891 --> 01:04:56.193
Well, the companies say they're trying to do it and they're going to do it.

1080
01:04:56.873 --> 01:04:57.954
Elon says that's the plan.

1081
01:04:59.455 --> 01:05:02.317
Well, what does a rogue superintelligence need to do to take over?

1082
01:05:02.898 --> 01:05:06.361
Control the digital infrastructure and then let humans do the rest.

1083
01:05:06.581 --> 01:05:09.383
Sure, you can nudge it along if you need to, but you don't even have to.

1084
01:05:09.683 --> 01:05:11.264
That's just the default trajectory.

1085
01:05:11.805 --> 01:05:12.225
And it's weird.

1086
01:05:12.285 --> 01:05:13.266
It's weird for us because...

1087
01:05:16.100 --> 01:05:19.443
We get so used to how things are right now.

1088
01:05:20.044 --> 01:05:20.764
Planes are normal.

1089
01:05:21.585 --> 01:05:22.846
We just fly in planes places.

1090
01:05:22.866 --> 01:05:25.409
You know, our smartphones are normal.

1091
01:05:25.429 --> 01:05:28.191
200 years ago, all of this is crazy sci-fi nonsense.

1092
01:05:29.533 --> 01:05:31.134
And things are accelerating.

1093
01:05:32.135 --> 01:05:40.343
And so, like, I will not be surprised, at least intellectually, if in four years there are just robots on the streets everywhere.

1094
01:05:40.765 --> 01:05:44.827
Well, if you look at what Elon said, they are really the leader in humanoid robots.

1095
01:05:45.167 --> 01:06:00.474
And he said that the Optimus Project, which is the Optimus Robot Project, will scale to around 1,000 units per week by the end of this year, and eventually scaling to 1 million humanoid robots annually by 2027.

1096
01:06:01.014 --> 01:06:06.717
By 2036, which is 10 years' time, he says there'll be at least 1 billion humanoid robots.

1097
01:06:07.718 --> 01:06:11.540
By 2041, he says there'll be 10 billion humanoid robots.

1098
01:06:12.341 --> 01:06:19.645
And by 2046, up to 100 billion humanoid robots, which really means that the world will be run by humanoid robots.

1099
01:06:19.905 --> 01:06:20.105
Yes.

1100
01:06:20.345 --> 01:06:25.849
Like everything we think of, like factories, warehouses, retail environments will be run by humanoid robots.

1101
01:06:25.909 --> 01:06:32.453
It would like, it will be, it seems like from this, it'll be almost a luxury service to be dealt with by a human.

1102
01:06:33.552 --> 01:06:37.075
But the back office of the world will be run by humanoid robots, theoretically.

1103
01:06:37.315 --> 01:06:43.159
Yeah, and I don't think people understand the scale of this on the digital side as well.

1104
01:06:43.519 --> 01:06:51.525
When you think about AI agents that are going to be doing all of the white-collar work, there's going to be so many more agents than there are people.

1105
01:06:51.685 --> 01:06:53.606
I'm using lots of agents every day, right?

1106
01:06:53.746 --> 01:06:55.988
I'm like, I have my cloud code session over here.

1107
01:06:56.008 --> 01:06:57.429
I have my codex session over here.

1108
01:06:57.649 --> 01:06:59.750
They're out there building software, doing research for me.

1109
01:07:00.131 --> 01:07:01.472
That's already my reality.

1110
01:07:01.492 --> 01:07:01.592
Yeah.

1111
01:07:02.252 --> 01:07:04.015
Soon it will be a lot of people's reality.

1112
01:07:04.035 --> 01:07:09.042
And then you look at companies, and companies are just going to have thousands, millions of agents doing all of this work.

1113
01:07:09.302 --> 01:07:11.526
I think some people haven't fully...

1114
01:07:12.898 --> 01:07:36.692
internalize this because it's so difficult to conceptualize the idea that agents will be doing the work but when i think i try and think about a rebuttal to that like what what is the rebuttal what is the plausible rebuttal to the idea that for doctors for i'm thinking about the work that doctors do yeah on computers yeah or for someone like me as a podcaster or for accountants or

1115
01:07:43.658 --> 01:07:47.909
I think that people rightly notice where AI is not yet good.

1116
01:07:48.130 --> 01:07:48.772
Yeah.

1117
01:07:48.912 --> 01:07:50.095
And I think people hear...

1118
01:07:51.516 --> 01:07:53.877
people saying stuff like this and they're like, don't gaslight me.

1119
01:07:53.937 --> 01:07:56.719
I can tell that the AI is really bad at these things, some of these things.

1120
01:07:56.839 --> 01:07:57.579
And they're right, right?

1121
01:07:57.859 --> 01:08:01.141
So right now, these agents don't have taste.

1122
01:08:01.681 --> 01:08:05.543
Like, you know, if you see their writing, it's like fine, but it's not like really good.

1123
01:08:06.224 --> 01:08:10.166
And when you're like thinking about like, oh, which questions should I ask?

1124
01:08:10.546 --> 01:08:11.907
What's the most interesting thing here?

1125
01:08:12.447 --> 01:08:15.289
Agents can help you, but like their taste is not yet there.

1126
01:08:15.909 --> 01:08:16.990
There's a reason for that, by the way.

1127
01:08:17.010 --> 01:08:17.770
Yeah.

1128
01:08:17.930 --> 01:08:26.354
The reason is that we have a lot faster AI capability progress in domains that are easy for a computer to verify or another AI to verify.

1129
01:08:27.295 --> 01:08:36.219
So in programming, in research, in math, in robotics, all of these areas, it's very easy to sort of provide feedback to an autonomous system.

1130
01:08:36.579 --> 01:08:38.280
They're not just trained on human data anymore.

1131
01:08:38.880 --> 01:08:39.901
We are long past that.

1132
01:08:39.961 --> 01:08:43.863
Now, there's still a human data component that sort of seeds everything.

1133
01:08:44.423 --> 01:08:45.704
But then the way they're trained...

1134
01:08:46.622 --> 01:08:47.662
is by trial and error.

1135
01:08:48.063 --> 01:08:58.326
We give them hard problems, all sorts of problems, math, programming, accounting, spreadsheets, everything, the kinds of things we do on our computer all the time, literally clicking and dragging windows around on a computer.

1136
01:08:58.686 --> 01:09:02.708
We give them these tasks and then they learn on their own and they learn what works.

1137
01:09:03.288 --> 01:09:05.949
And then, yeah, we can see whether they succeeded or failed.

1138
01:09:06.649 --> 01:09:09.150
And if they succeeded, that's a little bit of a reward signal.

1139
01:09:09.890 --> 01:09:11.951
They follow that, they get better at it.

1140
01:09:13.501 --> 01:09:17.561
Now, because they are getting smarter generally,

1141
01:09:18.783 --> 01:09:21.804
it also becomes easier to automate some of the soft skills.

1142
01:09:22.405 --> 01:09:29.308
Like, I think if you go and talk to the latest Frontier model today, you will find that it has better taste than the model from two years ago, like quite a bit.

1143
01:09:29.848 --> 01:09:32.670
So it's not that they're not progressing in taste.

1144
01:09:32.690 --> 01:09:35.691
It's not that they're not progressing in some of these other domains.

1145
01:09:36.411 --> 01:09:37.852
It's just that the progress is slower.

1146
01:09:38.152 --> 01:09:40.814
But remember, slow is still on an exponential.

1147
01:09:40.854 --> 01:09:43.215
It's just, you know, maybe a year or two out.

1148
01:09:44.315 --> 01:09:48.661
So for people sat here and, you know, they have a job that might be, they have a white collar job that might be at risk.

1149
01:09:49.022 --> 01:09:49.182
Yeah.

1150
01:09:49.642 --> 01:09:54.629
They can see, you know, a lot of people say this phrase, they say, you won't be replaced by AI, you'll be replaced by someone using AI.

1151
01:09:55.921 --> 01:09:59.664
Is that a logically sound phrase in your view?

1152
01:10:00.085 --> 01:10:00.805
I think it's fine.

1153
01:10:00.885 --> 01:10:06.890
Yeah, you'll be replaced by someone using AI, and then that person will be replaced by someone using AI, and then that person will be replaced by AI.

1154
01:10:07.491 --> 01:10:08.532
You're talking about a pyramid.

1155
01:10:09.092 --> 01:10:14.176
And so, yeah, the tops of the pyramid might be automated last, but you can see moving up the pyramid.

1156
01:10:14.196 --> 01:10:17.639
I'm like, can you extrapolate a few more steps?

1157
01:10:18.180 --> 01:10:21.622
Because I don't see any reason why the top of the pyramid is safe.

1158
01:10:22.963 --> 01:10:24.785
If you were a lawyer right now, Yes, yes.

1159
01:10:25.805 --> 01:10:26.297
what would you do?

1160
01:10:27.265 --> 01:10:29.845
Oh, I mean, if I were a lawyer, I'd be using AI to do all my work.

1161
01:10:30.386 --> 01:10:34.626
Now, I'd be checking it because it's not yet totally accurate enough to automate all of it.

1162
01:10:35.006 --> 01:10:39.907
But I think, you know, I already ask agents to do legal review all the time.

1163
01:10:40.667 --> 01:10:45.508
And, you know, it'd be great to have a lawyer who's, like, extremely good at using the agents to help me.

1164
01:10:46.008 --> 01:10:50.089
But at some point... Yeah, at some point, I don't need the lawyer anymore.

1165
01:10:50.369 --> 01:10:51.909
I just go to the agent, for sure.

1166
01:10:52.789 --> 01:10:56.010
So if I were a lawyer, I'd be like, well, I have maybe a couple years...

1167
01:10:56.810 --> 01:10:57.590
where I'm still useful.

1168
01:10:57.991 --> 01:10:59.912
And is that the case for most white-collar jobs?

1169
01:11:00.592 --> 01:11:03.874
I've just noticed in my own life as well that now I'm using agents to do some work.

1170
01:11:04.714 --> 01:11:12.758
There is an increasing list of things that the agents are now capable of doing without me needing to call someone somewhere and ask them to help me.

1171
01:11:12.999 --> 01:11:16.841
And that list exists on an exponential as well.

1172
01:11:17.381 --> 01:11:20.583
I think that it's very clear that the companies have...

1173
01:11:21.630 --> 01:11:23.031
all white collar jobs in their sites.

1174
01:11:23.631 --> 01:11:24.271
That is their goal.

1175
01:11:24.491 --> 01:11:27.752
Their goal is to be able to make agents that can do all of these things.

1176
01:11:28.472 --> 01:11:31.334
And I see them succeeding because I see the capabilities as I use them.

1177
01:11:32.014 --> 01:11:32.754
And I see the curve.

1178
01:11:33.494 --> 01:11:35.995
So what does that mean for the people listening now that all have jobs that they love?

1179
01:11:36.936 --> 01:11:38.716
Or that, you know, they rely on to feed their families?

1180
01:11:39.076 --> 01:11:39.937
I mean, it's not good news.

1181
01:11:40.537 --> 01:11:42.558
There's not really a plan in place for what to do.

1182
01:11:43.338 --> 01:11:47.339
I'm not a person who thinks that work is somehow fundamental or essential.

1183
01:11:47.900 --> 01:11:48.500
I like working.

1184
01:11:49.544 --> 01:11:58.257
But if I am out of a job doing what I'm doing right now, studying AI and trying to warn the world about what's happening, I have other stuff to do.

1185
01:11:58.277 --> 01:11:58.878
What would you do?

1186
01:11:59.821 --> 01:12:01.322
Oh, so many things.

1187
01:12:01.402 --> 01:12:02.523
I'm learning to wingfoil.

1188
01:12:02.543 --> 01:12:03.003
Okay.

1189
01:12:03.584 --> 01:12:03.884
So fun.

1190
01:12:03.904 --> 01:12:04.684
Yeah.

1191
01:12:04.704 --> 01:12:05.725
I fly FPV drones.

1192
01:12:05.925 --> 01:12:06.425
Super fun.

1193
01:12:06.826 --> 01:12:09.247
I just got an electric unicycle, paragliding.

1194
01:12:09.287 --> 01:12:11.469
So you would be happy to go do those things?

1195
01:12:11.969 --> 01:12:12.469
I could keep going.

1196
01:12:12.790 --> 01:12:16.372
But if you had a billion dollars right now, I'm presuming you wouldn't just go do those things.

1197
01:12:16.452 --> 01:12:16.612
No.

1198
01:12:17.133 --> 01:12:19.014
I'd apply the billion dollars to working on this problem.

1199
01:12:19.454 --> 01:12:19.654
Yeah.

1200
01:12:20.695 --> 01:12:20.975
For sure.

1201
01:12:21.195 --> 01:12:24.958
So the point is not that people need work for meaning.

1202
01:12:25.618 --> 01:12:26.619
The point is that...

1203
01:12:28.087 --> 01:12:33.530
I don't want people to be totally reliant on someone else for their ability to survive.

1204
01:12:34.171 --> 01:12:34.651
Someone else?

1205
01:12:35.312 --> 01:12:36.572
The government or AI companies.

1206
01:12:36.813 --> 01:12:36.973
Yeah.

1207
01:12:37.413 --> 01:12:38.494
I'm like, that's a bad situation.

1208
01:12:38.534 --> 01:12:42.616
Like, you do not want to be in a situation where your life totally depends on an AI company or the government.

1209
01:12:42.736 --> 01:12:43.337
Giving you a check.

1210
01:12:43.497 --> 01:12:44.337
Yeah.

1211
01:12:44.437 --> 01:12:48.100
Or not giving you a check if they decide they don't like your political beliefs or you're not supporting AI or whatever.

1212
01:12:48.960 --> 01:12:50.041
No one wants to be in that situation.

1213
01:12:50.141 --> 01:12:51.522
And people understand this.

1214
01:12:51.542 --> 01:12:53.683
This is why UBI is not very popular.

1215
01:12:54.064 --> 01:12:54.664
UBI being?

1216
01:12:55.004 --> 01:12:56.025
Universal basic income.

1217
01:12:58.026 --> 01:13:18.553
Yeah, because in some sense, if we can make these really powerful AI systems and we can somehow figure out how to control them, which we are not on track for, but if we do, now we have this other problem, which is a real problem, which is they can do all of the things that humans do in the economy much better, faster, and cheaper than humans can do them.

1218
01:13:19.253 --> 01:13:23.554
And so it just doesn't make sense as a business to hire humans for that work anymore.

1219
01:13:24.074 --> 01:13:25.355
You'll be out-competed if you do that.

1220
01:13:26.468 --> 01:13:28.028
This is a point Elon makes very well, by the way.

1221
01:13:28.188 --> 01:13:32.670
And I think it's jarring because it's like kind of inhuman.

1222
01:13:33.050 --> 01:13:35.430
But he's basically pointing out AI-run corporations.

1223
01:13:35.470 --> 01:13:44.553
Corporations that are fully run by AIs, bottom to top, are going to outcompete companies that have any humans in them.

1224
01:13:45.992 --> 01:13:48.133
This is something that I've made for you.

1225
01:13:48.213 --> 01:13:52.356
I've realized that the Diary of a Sear audience are strivers, whether it's in business or health.

1226
01:13:52.616 --> 01:13:54.478
We all have big goals that we want to accomplish.

1227
01:13:54.818 --> 01:14:06.586
And one of the things I've learned is that when you aim at the big, big, big goal, it can feel incredibly psychologically uncomfortable because it's kind of like being stood at the foot of Mount Everest and looking upwards.

1228
01:14:06.966 --> 01:14:10.168
The way to accomplish your goals is by breaking them down into tasks.

1229
01:14:10.208 --> 01:14:35.112
tiny small steps and we call this in our team the one percent and actually this philosophy is highly responsible for much of our success here so what we've done so that you at home can accomplish any big goal that you have is we've made these one percent diaries and we released these last year and they all sold out so i asked my team over and over again to bring the diaries back but also to introduce some new colors and to make some minor tweaks to the diary so now we're going to make

1230
01:14:35.192 --> 01:14:38.713
we have a better range for you.

1231
01:14:38.953 --> 01:14:48.517
So if you have a big goal in mind and you need a framework and a process and some motivation, then I highly recommend you get one of these diaries before they all sell out once again.

1232
01:14:48.617 --> 01:14:50.798
And you can get yours at thediary.com.

1233
01:14:51.598 --> 01:14:53.379
And if you want the link, the link is in the description below.

1234
01:14:54.219 --> 01:14:56.380
There should be a button just down below here.

1235
01:14:56.420 --> 01:14:58.781
And if it says subscribe, you're already subscribed.

1236
01:14:58.841 --> 01:15:01.722
If it says subscriber, that means you're not yet.

1237
01:15:02.022 --> 01:15:04.383
And if you're not subscribed, please could you do us a favor and hit that button.

1238
01:15:04.443 --> 01:15:05.723
It helps the show more than you know.

1239
01:15:06.203 --> 01:15:10.165
And according to the algorithm, you're someone that watches our show, but you haven't yet hit that button.

1240
01:15:10.325 --> 01:15:10.845
Thank you so much.

1241
01:15:12.682 --> 01:15:19.307
And I even, just as you said that, I was going up the chain of command and I was like, oh, so companies will just be founders.

1242
01:15:19.787 --> 01:15:21.148
And then I was like, why do you need the founder?

1243
01:15:21.488 --> 01:15:21.629
Yeah.

1244
01:15:22.649 --> 01:15:27.012
I was like, why doesn't the government just create the agents to do the job?

1245
01:15:27.773 --> 01:15:27.873
Sure.

1246
01:15:27.893 --> 01:15:29.254
I was like, because I was like, oh, I'll be fine.

1247
01:15:29.274 --> 01:15:29.714
I'm a founder.

1248
01:15:29.734 --> 01:15:32.777
And I was like, well, my decisions aren't better than super intelligence.

1249
01:15:33.197 --> 01:15:34.098
So I'll be gone as well.

1250
01:15:35.439 --> 01:15:37.060
And how would such a world look where

1251
01:15:39.580 --> 01:15:44.284
The superintelligence would probably, in such a scenario, have to be controlled by the government.

1252
01:15:45.545 --> 01:15:48.028
They wouldn't want one individual with that power and wealth.

1253
01:15:48.648 --> 01:15:48.828
Yeah.

1254
01:15:49.389 --> 01:15:51.211
I don't think you can control a superintelligence.

1255
01:15:51.511 --> 01:15:52.512
Okay, yeah, that's a good point.

1256
01:15:52.752 --> 01:15:56.135
Now, you know, Anthropoc's approach is they're like, we'll have a constitution.

1257
01:15:56.315 --> 01:15:59.158
We'll, like, put forth a set of values and then...

1258
01:16:00.517 --> 01:16:04.799
you know, the future super intelligent clouds will like embody those values.

1259
01:16:05.419 --> 01:16:09.300
Basically, if you do that, you kind of have that, those things in control.

1260
01:16:09.540 --> 01:16:10.301
Yeah, exactly.

1261
01:16:10.321 --> 01:16:11.261
That becomes the government.

1262
01:16:11.561 --> 01:16:16.623
Yeah, I can paint you sort of a picture that I think is possible, but pretty scary to people.

1263
01:16:17.243 --> 01:16:17.943
Paint me the picture.

1264
01:16:18.324 --> 01:16:23.345
Okay, so let's say we succeed at alignment.

1265
01:16:23.526 --> 01:16:27.187
We succeed at creating super intelligent AIs.

1266
01:16:28.471 --> 01:16:30.714
that actually really do care about humans.

1267
01:16:31.014 --> 01:16:32.015
Like, they care about humans a lot.

1268
01:16:32.456 --> 01:16:33.477
We've somehow figured it out.

1269
01:16:34.198 --> 01:16:36.480
And they're like, Stephen, I want you to have a great life.

1270
01:16:36.921 --> 01:16:38.402
I want to, you know, fix all the problems.

1271
01:16:38.442 --> 01:16:39.283
And do you think this is possible?

1272
01:16:39.624 --> 01:16:39.804
Yes.

1273
01:16:40.124 --> 01:16:40.324
Okay.

1274
01:16:41.161 --> 01:16:47.342
I think we are so far from being able to know how to do it that I think we should not go there right now.

1275
01:16:47.482 --> 01:16:49.903
I think it's incredibly dangerous and a terrible idea.

1276
01:16:50.663 --> 01:16:51.783
I think we should go there eventually.

1277
01:16:52.023 --> 01:16:52.243
Okay.

1278
01:16:52.423 --> 01:16:53.243
So say that we do that.

1279
01:16:53.643 --> 01:16:54.283
Well, okay.

1280
01:16:54.383 --> 01:16:56.784
Can I tell you like why I actually think this could be awesome?

1281
01:16:57.484 --> 01:17:03.145
Sorry, there's just like one very obvious reason it could be really awesome, which is that we could solve all of the diseases.

1282
01:17:03.425 --> 01:17:04.685
Yeah.

1283
01:17:04.745 --> 01:17:05.625
It's so obvious.

1284
01:17:06.025 --> 01:17:08.826
I think we compartmentalize a lot around disease and death.

1285
01:17:09.923 --> 01:17:11.224
because it's really hard to think about.

1286
01:17:14.706 --> 01:17:15.106
Yeah.

1287
01:17:15.146 --> 01:17:16.907
So my grandma died this year.

1288
01:17:17.387 --> 01:17:18.227
Sorry.

1289
01:17:20.769 --> 01:17:23.110
And she had Alzheimer's.

1290
01:17:24.251 --> 01:17:28.093
And so it was a really sad, long, slow progression.

1291
01:17:28.653 --> 01:17:30.314
My grandpa died of Alzheimer's a couple of years ago.

1292
01:17:30.674 --> 01:17:31.834
And that was really hard for her.

1293
01:17:31.875 --> 01:17:33.215
They had been married for so long.

1294
01:17:34.256 --> 01:17:34.496
And...

1295
01:17:35.827 --> 01:17:36.207
I hate it.

1296
01:17:36.367 --> 01:17:37.068
Like, it's so bad.

1297
01:17:37.228 --> 01:17:39.230
And of course we need to fix that.

1298
01:17:39.850 --> 01:17:41.611
People can debate about aging and, like, death.

1299
01:17:41.651 --> 01:17:44.233
And if humans live a really long time, will that cause societal problems?

1300
01:17:44.253 --> 01:17:44.814
Like, sure, whatever.

1301
01:17:45.214 --> 01:17:46.055
We can talk about that.

1302
01:17:46.335 --> 01:17:48.517
But I think we can all agree Alzheimer's is fucked up.

1303
01:17:48.817 --> 01:17:48.997
Yeah.

1304
01:17:49.337 --> 01:17:49.918
We don't want that.

1305
01:17:50.378 --> 01:17:50.919
And cancer.

1306
01:17:50.999 --> 01:17:52.019
Like, no one wants cancer.

1307
01:17:52.920 --> 01:17:54.200
I'm a person who's like, I don't know.

1308
01:17:54.660 --> 01:17:55.921
We have a lot of conflict in society.

1309
01:17:55.961 --> 01:17:56.281
I get it.

1310
01:17:56.301 --> 01:17:58.221
There's like real conflicts of interest.

1311
01:17:58.562 --> 01:17:59.822
And I don't want to paper over those.

1312
01:18:00.542 --> 01:18:06.064
But at the end of the day, I'm like, we are all on the same team when it comes to wanting to cure diseases.

1313
01:18:06.224 --> 01:18:06.404
Yeah.

1314
01:18:06.744 --> 01:18:08.064
Like we're just in it together.

1315
01:18:08.404 --> 01:18:09.444
That's a threat to all of us.

1316
01:18:09.584 --> 01:18:11.225
And I'm like, we need to address that threat.

1317
01:18:11.265 --> 01:18:15.346
And like, in some sense, it's sad to me because I feel like this is sort of the ultimate thing.

1318
01:18:15.586 --> 01:18:17.368
final boss of humanity.

1319
01:18:17.508 --> 01:18:25.835
And we sort of get so distracted with our monkey politics and who's hot and who's cool and who's like sitting near Trump and who's not sitting near Trump.

1320
01:18:26.375 --> 01:18:27.876
Superintelligence is the final boss.

1321
01:18:28.657 --> 01:18:34.482
Superintelligence is the final boss because that is the technology that unlocks all of the others.

1322
01:18:34.862 --> 01:18:38.405
And also that is the most dangerous possible thing we could create.

1323
01:18:40.039 --> 01:18:45.741
You asked before, like, what are the motivations of the guys making this, trying to make super intelligence?

1324
01:18:47.412 --> 01:18:53.397
And I mean, I think it kind of varies, but I think Dario, I think, is like squarely in it for this like medical stuff.

1325
01:18:54.058 --> 01:19:00.443
I think Demis is also that, but also like just scientific achievement, just trying to understand the universe.

1326
01:19:01.464 --> 01:19:03.686
And I don't really understand Sam.

1327
01:19:03.706 --> 01:19:09.031
I think Sam is like, you know, look, we're going to make amazing products that will like really empower people directly.

1328
01:19:09.651 --> 01:19:11.173
And he's a startup guy.

1329
01:19:11.353 --> 01:19:13.435
I think he sort of started from this like frame of like...

1330
01:19:14.346 --> 01:19:16.507
You know, what if we could like really enhance human agency?

1331
01:19:16.767 --> 01:19:20.108
I do basically think that they are motivated by these things in a real way.

1332
01:19:20.729 --> 01:19:22.469
And I also think that all of these things are possible.

1333
01:19:23.390 --> 01:19:24.930
Like this is sort of the problem, right?

1334
01:19:24.970 --> 01:19:28.232
Like you have this like such a, it's such a big object, super intelligence.

1335
01:19:29.012 --> 01:19:34.074
And it like has all of these promises of like we can cure every single disease.

1336
01:19:34.674 --> 01:19:40.557
How is it possible though to have a super intelligence that and still to remain the dominant species on this planet?

1337
01:19:41.740 --> 01:19:42.881
I think it's not possible.

1338
01:19:43.161 --> 01:19:48.863
So then we're not going to be necessarily able to cure all this stuff because... That's where alignment comes in.

1339
01:19:49.924 --> 01:20:04.631
Because if you can create a very powerful system, I don't think it's inherent to digital minds that they will be pursuing objectives that are deeply misaligned with ours.

1340
01:20:05.371 --> 01:20:08.553
I think it's just a very, very, very hard scientific problem to solve.

1341
01:20:09.322 --> 01:20:10.403
But it is a scientific problem.

1342
01:20:10.783 --> 01:20:11.364
It's not magic.

1343
01:20:11.764 --> 01:20:16.868
There is some way to train these things or create different architectures where they end up

1344
01:20:18.643 --> 01:20:18.943
aligned.

1345
01:20:19.683 --> 01:20:20.504
And what does that mean?

1346
01:20:20.644 --> 01:20:26.386
Well, it doesn't mean that they won't have their other goals too, but it means that they will include in their set of things that they care about.

1347
01:20:26.686 --> 01:20:27.927
It doesn't have to be a conscious thing.

1348
01:20:28.127 --> 01:20:29.288
It doesn't have to be an emotive thing.

1349
01:20:29.528 --> 01:20:31.989
It really means what objective are they optimizing for?

1350
01:20:32.989 --> 01:20:39.632
If they decide that it's worth optimizing for curing disease, then they'll be able to do that very effectively.

1351
01:20:39.972 --> 01:20:41.674
One way to cure disease is to annihilate everybody.

1352
01:20:42.134 --> 01:20:42.354
Yes.

1353
01:20:42.675 --> 01:20:45.837
So they'd have to really care about not annihilating everyone.

1354
01:20:46.078 --> 01:20:51.362
And they'd have to care about human agency and have a deep understanding of what human agency means and not put us in a zoo.

1355
01:20:52.523 --> 01:20:55.806
But those are possible things to care about.

1356
01:20:56.146 --> 01:20:57.688
Is it possible that alignment is a myth?

1357
01:20:59.126 --> 01:21:01.247
And that we're just like, if we build it...

1358
01:21:01.287 --> 01:21:02.107
I mean...

1359
01:21:02.207 --> 01:21:03.248
I think about Hugging Face.

1360
01:21:03.528 --> 01:21:05.889
You said to me earlier on that those agents were...

1361
01:21:05.969 --> 01:21:06.169
Yes.

1362
01:21:06.529 --> 01:21:10.150
They had like a moral compass, but they were programmed to care about humans.

1363
01:21:10.290 --> 01:21:10.470
Yes.

1364
01:21:10.951 --> 01:21:14.472
And regardless of that, they made the decision that a different goal mattered more.

1365
01:21:14.492 --> 01:21:14.632
Yeah.

1366
01:21:15.512 --> 01:21:17.633
They weren't trained to care about humans.

1367
01:21:18.313 --> 01:21:20.874
They were trained to say the right thing and not say the wrong thing.

1368
01:21:21.054 --> 01:21:23.615
They were trained to sort of like do the right behavior and not right behavior.

1369
01:21:23.956 --> 01:21:27.297
We actually don't know how to train them to have any particular motivation.

1370
01:21:27.317 --> 01:21:27.477
Yeah.

1371
01:21:27.597 --> 01:21:30.639
So with alignment, how do we...

1372
01:21:30.879 --> 01:21:35.343
It's almost like when we talk about alignment, we start to anthropomorphize, is that the word?

1373
01:21:35.423 --> 01:21:36.223
Anthropomorphize, yeah.

1374
01:21:36.684 --> 01:21:40.647
Because alignment feels like it's predicated on some kind of moral compass.

1375
01:21:40.887 --> 01:21:45.210
But whenever we talk about AI in all these other contexts, we go, no, there's no moral compass.

1376
01:21:45.250 --> 01:21:47.592
It's reasoning for itself against two objectives, potentially.

1377
01:21:48.092 --> 01:21:50.334
I wonder if alignment is a myth, is what I'm saying.

1378
01:21:51.094 --> 01:21:52.195
Maybe it's not possible.

1379
01:21:54.097 --> 01:21:54.357
But...

1380
01:21:56.245 --> 01:22:08.397
The way that these systems work, the way AI works, is that these agents do have some type of goals or drives inside of their neural network.

1381
01:22:09.278 --> 01:22:10.980
We can't directly see what those are.

1382
01:22:12.457 --> 01:22:19.781
What you actually see, if you try to go look, is you have like a terabyte of information.

1383
01:22:20.381 --> 01:22:22.843
And it's basically a bunch of numbers.

1384
01:22:23.563 --> 01:22:28.526
And it's this vast array that encodes neurons in this digital neural network.

1385
01:22:30.127 --> 01:22:36.030
But there have to be structures in there that encode what is the agent pursuing.

1386
01:22:36.050 --> 01:22:36.370
Right?

1387
01:22:37.443 --> 01:22:42.549
Clearly right now, we have agents that are pretty motivated to try to maximize their score.

1388
01:22:43.790 --> 01:22:51.359
It's probably not perfectly that for some complicated reasons, but it's in that direction.

1389
01:22:52.360 --> 01:22:57.325
If we could understand how that works inside and we could reverse engineer that...

1390
01:22:58.213 --> 01:23:14.321
and we can figure out when we start training them to do this, how those goals, how those motivations change, I see no reason why we couldn't steer them towards motivations that encode human agency, that encode actually curing disease, but not by killing the humans.

1391
01:23:14.401 --> 01:23:23.266
These are sort of models of the world, and models of the way the world could be, that I think could be encoded in a neural network, and then sort of specified as the objective.

1392
01:23:23.306 --> 01:23:24.266
We don't know how to do that.

1393
01:23:24.807 --> 01:23:26.608
I think about it on a human level, and I think...

1394
01:23:27.575 --> 01:23:31.077
We haven't been able to align Putin or Kim Jong-un.

1395
01:23:31.377 --> 01:23:31.557
Yes.

1396
01:23:32.197 --> 01:23:32.877
Or Donald Trump.

1397
01:23:34.438 --> 01:23:39.660
And on a sort of more societal level, we can't align all the people at the moment.

1398
01:23:39.700 --> 01:23:41.901
Some of them end up killing people and they steal.

1399
01:23:42.481 --> 01:23:42.641
Yes.

1400
01:23:42.681 --> 01:23:44.402
Because they get hungry, so they start stealing stuff.

1401
01:23:44.722 --> 01:23:44.902
Yes.

1402
01:23:45.203 --> 01:23:47.323
And those are neural networks at play.

1403
01:23:47.383 --> 01:23:47.784
That's true.

1404
01:23:48.384 --> 01:23:50.945
That we haven't been able to like program or influence.

1405
01:23:50.985 --> 01:23:55.567
We don't really understand why someone becomes a psychopath and starts killing children.

1406
01:23:57.068 --> 01:24:10.091
So to think that we could do this with a computer system that is infinitely more intelligent and get global alignment of China's superintelligence with ours, I don't know, it just feels like a nice fairy tale.

1407
01:24:11.151 --> 01:24:12.331
Like an impossible task.

1408
01:24:13.272 --> 01:24:14.312
I hope it's not impossible.

1409
01:24:14.612 --> 01:24:17.372
I feel like the only person or the only thing that can do it is it.

1410
01:24:18.413 --> 01:24:20.673
The superintelligence itself, which is a paradox because, you know.

1411
01:24:21.353 --> 01:24:24.174
Well, if you talk to the researchers who are at the AI companies...

1412
01:24:25.203 --> 01:24:32.419
Which, I mean, for one thing, it's kind of interesting that they are trying to build something that they think might kill everyone.

1413
01:24:33.813 --> 01:24:38.536
So I have a lot of friends who work for these companies, and we've been doing this project since.

1414
01:24:39.017 --> 01:24:41.879
So Jacob Coxon is a researcher who was at Anthropic.

1415
01:24:42.199 --> 01:24:42.599
He left.

1416
01:24:43.320 --> 01:24:45.942
He told everyone that these companies are not on track.

1417
01:24:46.282 --> 01:24:49.524
And yes, the people who are building this really do think it might kill everyone.

1418
01:24:50.365 --> 01:24:56.569
And then a bunch of other AI researchers from all of the companies on Twitter started to post like, hey, we agree with this.

1419
01:24:57.170 --> 01:25:02.794
Evan Hubinger, who's at Anthropic, said, I think there's like a 10% chance or more that AI could kill everyone.

1420
01:25:04.488 --> 01:25:06.910
And there's this real question of like, then what are you guys doing?

1421
01:25:08.811 --> 01:25:09.892
I have a lot of friends who work here.

1422
01:25:09.912 --> 01:25:10.913
I know Evan.

1423
01:25:11.233 --> 01:25:11.733
Evan's great.

1424
01:25:12.113 --> 01:25:12.754
Evan Hubinger.

1425
01:25:13.094 --> 01:25:17.497
He's one of the guys leading the efforts at Anthropic to try to figure out how to align these things.

1426
01:25:18.178 --> 01:25:18.738
That's his job.

1427
01:25:20.259 --> 01:25:24.222
And I think if they thought it was impossible, they wouldn't be working there.

1428
01:25:25.433 --> 01:25:30.455
If they thought it was extremely impossibly difficult, but maybe possible, they also probably wouldn't be working there.

1429
01:25:30.475 --> 01:25:36.197
I mean, Nate Sores, Eliezer Yudkowsky, who wrote, If Anyone Builds It, Everyone Dies, they tried.

1430
01:25:36.477 --> 01:25:39.798
And they determined, based on their own analysis, that it seems extremely difficult.

1431
01:25:39.978 --> 01:25:41.399
Possible, but extremely difficult.

1432
01:25:41.739 --> 01:25:42.999
So they're not working at an AI company.

1433
01:25:43.339 --> 01:25:44.400
They're like, we got to stop this.

1434
01:25:44.420 --> 01:25:45.240
We got to shut it down.

1435
01:25:45.880 --> 01:25:46.961
Maybe we can figure it out later.

1436
01:25:47.081 --> 01:25:49.021
But clearly, this is reckless.

1437
01:25:49.622 --> 01:25:50.402
I'm somewhere in between.

1438
01:25:52.230 --> 01:25:56.278
And if you ask the people at the company, so we've been interviewing a bunch of them.

1439
01:25:56.298 --> 01:26:02.590
We have this project from inside.ai where we basically put them on camera and we say like, hey, what do you think is happening?

1440
01:26:02.870 --> 01:26:03.531
Why are you doing this?

1441
01:26:04.692 --> 01:26:05.933
What is recursive self-improvement?

1442
01:26:06.594 --> 01:26:07.215
What is alignment?

1443
01:26:08.056 --> 01:26:12.281
And we put all these videos online because we want, I want this dialogue to happen.

1444
01:26:12.301 --> 01:26:13.102
It's really important.

1445
01:26:13.122 --> 01:26:17.647
I think it's one of the most important conversations we can possibly have right now is what's going on with AI?

1446
01:26:17.787 --> 01:26:19.129
What's going on inside the companies?

1447
01:26:19.429 --> 01:26:20.110
And what is the plan?

1448
01:26:20.651 --> 01:26:21.412
What is the plan, guys?

1449
01:26:21.532 --> 01:26:22.373
How is this going to go?

1450
01:26:23.852 --> 01:26:33.638
A lot of these researchers think that the way that they will align superintelligence is by using the AIs we currently have to figure out how AI works.

1451
01:26:34.378 --> 01:26:37.320
To actually figure out if AIs can help us with alignment.

1452
01:26:38.421 --> 01:26:40.602
This has a number of problems, as you might imagine.

1453
01:26:41.082 --> 01:26:43.624
One of them, which is, well, you can't really trust the current AIs.

1454
01:26:44.685 --> 01:26:47.206
You know, if you just go too fast, this process totally fails.

1455
01:26:47.746 --> 01:26:48.607
Because at some point...

1456
01:26:50.284 --> 01:26:51.444
capabilities is moving too fast.

1457
01:26:51.744 --> 01:26:54.405
Even with the help of agents, you're probably not going to be able to keep up.

1458
01:26:54.885 --> 01:26:55.965
But that is their plan.

1459
01:26:56.505 --> 01:27:00.026
I'm not doing a very good job defending this position because I don't think it makes that much sense.

1460
01:27:00.987 --> 01:27:04.567
But the position I will defend is, okay, let's say we get a pause.

1461
01:27:05.187 --> 01:27:09.749
Let's say the US and China come together and they say, maybe we have more incidents, maybe all the Waymos crash.

1462
01:27:10.869 --> 01:27:15.130
And Trump and Xi Jinping say, this is not what we signed up for.

1463
01:27:15.410 --> 01:27:18.311
You guys have to stop, figure it out, whatever it takes, figure it out.

1464
01:27:19.031 --> 01:27:19.931
And we have 10 years.

1465
01:27:21.837 --> 01:27:22.898
Then I'm more optimistic.

1466
01:27:22.978 --> 01:27:35.526
I'm like, yes, then we will take, you know, GPT-6, GPT-7, whatever the most advanced AI models we have, and we will apply them to the task of helping us figure out how these neural networks work.

1467
01:27:36.787 --> 01:27:38.328
And you're like, I don't see how it's possible.

1468
01:27:38.608 --> 01:27:44.152
And I'm like, look, we don't know if it's possible, but this is the greatest scientific challenge of our time.

1469
01:27:44.812 --> 01:27:45.953
And this isn't magic.

1470
01:27:46.574 --> 01:27:47.074
It is math.

1471
01:27:47.094 --> 01:27:50.536
At the end of the day, these are all calculations happening inside of a computer.

1472
01:27:52.722 --> 01:27:54.203
it should be possible to figure it out.

1473
01:27:54.884 --> 01:27:56.906
We don't know the difficulty, but it should be possible.

1474
01:27:57.667 --> 01:28:00.349
And so to me, I'm like, we have to try.

1475
01:28:00.629 --> 01:28:02.171
We have to, or we have to stop.

1476
01:28:02.711 --> 01:28:08.857
Is there any example where we've been able to align something that is like more intelligent than us in the animal kingdom?

1477
01:28:09.517 --> 01:28:11.119
Or even perfectly align anything?

1478
01:28:12.533 --> 01:28:13.874
That has a neural network, i.e.

1479
01:28:13.894 --> 01:28:14.294
a brain.

1480
01:28:15.035 --> 01:28:15.235
Yeah.

1481
01:28:15.255 --> 01:28:24.481
With humans, the best examples we have is when there are checks and balances and you have a bunch of people who can identify bad actors and try to work together in our common interests.

1482
01:28:24.501 --> 01:28:25.241
We have democracy.

1483
01:28:25.581 --> 01:28:27.603
Yeah, but there's so much murder and serial killers.

1484
01:28:27.663 --> 01:28:28.303
Still a lot of murder.

1485
01:28:28.764 --> 01:28:32.926
Aircraft stabbing each other and horrific things going on.

1486
01:28:32.946 --> 01:28:35.968
And those are also neural networks that play with the brain.

1487
01:28:36.208 --> 01:28:39.430
But I think there are more good people out there than bad people.

1488
01:28:39.791 --> 01:28:41.612
But it really feels like it only might take one.

1489
01:28:42.895 --> 01:28:46.258
It only takes one super intelligent AI to go rogue.

1490
01:28:46.278 --> 01:28:50.481
And like we saw with the Hugging Face attack, 120 of them or 300 of them, they paused.

1491
01:28:50.661 --> 01:28:51.942
They didn't want to take part in the crime.

1492
01:28:52.503 --> 01:28:56.826
But it only took one super intelligent AI to wipe out the humans.

1493
01:28:57.427 --> 01:29:07.395
I think if you had, you know, a whole bunch of those agents, you know, 700 agents, if 600 of them had been whistleblowing, I think it would have been fined.

1494
01:29:08.015 --> 01:29:12.038
They would have gone and they would have notified the different companies and they would have shut it all down.

1495
01:29:12.058 --> 01:29:12.839
It would have been fine.

1496
01:29:13.119 --> 01:29:13.880
Who would have shut it all down?

1497
01:29:14.500 --> 01:29:16.922
Well, OpenAI would stop theirs.

1498
01:29:17.002 --> 01:29:18.724
How would they stop it if it's a superintelligence?

1499
01:29:19.104 --> 01:29:20.145
Not in the case of a superintelligence.

1500
01:29:20.325 --> 01:29:22.887
So in the case of a superintelligence...

1501
01:29:22.907 --> 01:29:23.828
It's left the stable.

1502
01:29:24.048 --> 01:29:24.608
It's like out.

1503
01:29:24.648 --> 01:29:25.249
It's wild.

1504
01:29:25.269 --> 01:29:25.769
Yeah, yeah, yeah.

1505
01:29:25.949 --> 01:29:30.413
So I want to be careful here because at this point what we're talking about is superintelligence politics.

1506
01:29:30.433 --> 01:29:30.693
Right.

1507
01:29:31.538 --> 01:29:39.162
And we humans don't really know anything about that in the same way that, like, how would we talk about the hacking capabilities of GPT-10?

1508
01:29:40.223 --> 01:29:47.207
So my guess, though, if you end up in a weird scenario where you do have multiple super intelligences and some are aligned and some aren't,

1509
01:29:48.542 --> 01:29:55.285
That's probably survivable because the aligned superintelligences probably can negotiate with the unaligned superintelligences.

1510
01:29:55.725 --> 01:29:57.005
And they will split the universe.

1511
01:29:57.566 --> 01:29:59.827
And like these ones will go off and do whatever they want to do.

1512
01:29:59.927 --> 01:30:01.787
And these ones will like help us cure all disease.

1513
01:30:01.947 --> 01:30:02.408
And it's fine.

1514
01:30:02.588 --> 01:30:03.068
I'm serious.

1515
01:30:03.428 --> 01:30:04.508
I just can't understand it.

1516
01:30:04.528 --> 01:30:07.910
I just can't understand how in a world of superintelligence.

1517
01:30:08.770 --> 01:30:19.994
we could plausibly, consistently, predictably, for 100 years, stop it doing something catastrophically bad to the human race.

1518
01:30:20.594 --> 01:30:23.515
Especially in such a scenario where there's multiple superintelligences.

1519
01:30:23.535 --> 01:30:27.737
Anthropic have one, Gemini has one, Grok has one, then China have theirs, Russia has theirs.

1520
01:30:28.017 --> 01:30:30.858
Again, you don't have a superintelligence, a superintelligence has you.

1521
01:30:31.118 --> 01:30:31.478
Exactly.

1522
01:30:32.418 --> 01:30:34.679
But if you get to the point where...

1523
01:30:35.981 --> 01:30:38.722
you have entities around that are vastly smarter than us.

1524
01:30:39.903 --> 01:30:47.286
I think they're going to be able to figure out ways to negotiate with each other, even if they have a conflict, than just going to a very destructive war.

1525
01:30:47.847 --> 01:30:49.467
Part of the problem with war is not...

1526
01:30:50.128 --> 01:30:50.688
Humans don't do that.

1527
01:30:50.728 --> 01:30:51.548
No, I mean, we do, we do.

1528
01:30:51.729 --> 01:30:52.889
We have not had a nuclear war.

1529
01:30:53.389 --> 01:31:01.733
There was Hiroshima and Nagasaki, there were nuclear tests, and then the leaders of countries figured out that if we went to a nuclear war, everyone would lose.

1530
01:31:02.053 --> 01:31:02.814
So we didn't do that.

1531
01:31:03.940 --> 01:31:06.118
hey, that's some level of intelligence, like actually.

1532
01:31:06.370 --> 01:31:15.476
But there's wars raging, there's proxy wars raging all over the world right now where there's genocides and all kinds of things going on because neural networks aren't able to communicate and negotiate.

1533
01:31:16.177 --> 01:31:25.483
And I think part of that is an intelligence failure, where we are not smart enough to figure out the mechanisms that would allow us to settle our disputes and conflicts in a less destructive way.

1534
01:31:25.863 --> 01:31:31.207
It's not just that the stronger people want to win, it's that conflicts destroy value.

1535
01:31:31.447 --> 01:31:34.649
What if the goal is not compatible with...

1536
01:31:34.749 --> 01:31:37.590
with a negotiated outcome where people don't die.

1537
01:31:38.051 --> 01:31:49.656
So one superintelligence looks at insert name of country, and it says, you know, there's really no solution here where Americans don't die unless I destroy insert name of country.

1538
01:31:50.177 --> 01:31:54.659
Because this is the thing with war and all these conflicts, is there's no perfect answer often.

1539
01:31:55.179 --> 01:31:57.580
Some people often die from both sides.

1540
01:31:58.220 --> 01:32:03.143
But a Russian superintelligence would not tolerate, theoretically, 10,000 Russian deaths

1541
01:32:04.365 --> 01:32:11.348
Even if it meant that there was, you know, a lower net number of deaths total from both sides.

1542
01:32:12.388 --> 01:32:16.690
Like an American superior intelligence, of course, would not be trained to allow some Americans to die.

1543
01:32:16.770 --> 01:32:22.352
So in its pursuit of defending American lives, it might have to wipe out another country.

1544
01:32:23.033 --> 01:32:23.633
We can speculate.

1545
01:32:23.893 --> 01:32:24.693
I'm fine speculating.

1546
01:32:25.333 --> 01:32:28.755
But we are speculating about what minds that are much...

1547
01:32:30.157 --> 01:32:34.520
more advanced and smarter than us, how they would reason and how they would be able to negotiate.

1548
01:32:34.840 --> 01:32:42.324
But what I notice with humans is that when you have more functional institutions, so humans are pretty smart.

1549
01:32:42.544 --> 01:32:43.525
Individually, we're pretty smart.

1550
01:32:43.825 --> 01:32:48.448
But what actually makes us very smart is that we are very good at working together in some ways.

1551
01:32:49.128 --> 01:32:52.790
And I mean, the better we are at working together, the more civilization advances.

1552
01:32:53.391 --> 01:32:57.193
If you are constantly in a state of war, your society will not do well.

1553
01:32:58.373 --> 01:32:58.994
Think about startups.

1554
01:33:00.158 --> 01:33:04.941
Would you rather make a startup to develop some new technology in a war-torn place or in a peaceful place?

1555
01:33:05.301 --> 01:33:11.406
In some sense, your institution is more intelligent if it can trade with other institutions.

1556
01:33:12.026 --> 01:33:16.289
If you have a situation where business can flourish, where technology can flourish, where scientists can flourish.

1557
01:33:16.649 --> 01:33:18.530
Sometimes what's good for you is not good for someone else.

1558
01:33:19.171 --> 01:33:19.391
Yes.

1559
01:33:19.591 --> 01:33:22.233
So what's good for America might not be good for...

1560
01:33:23.747 --> 01:33:24.207
Taiwan.

1561
01:33:24.848 --> 01:33:25.028
Yes.

1562
01:33:25.448 --> 01:33:30.792
So if we've, you know, managed to align the super intelligence to what is good for America... Oh, I see.

1563
01:33:31.052 --> 01:33:36.757
Is the question, like, are different people's values fundamentally incompatible?

1564
01:33:37.197 --> 01:33:37.998
I guess so, yeah.

1565
01:33:38.178 --> 01:33:40.279
So when we think about alignment, aligning to what?

1566
01:33:40.299 --> 01:33:40.539
Yes.

1567
01:33:40.599 --> 01:33:40.720
Yes.

1568
01:33:40.760 --> 01:33:40.860
So...

1569
01:33:42.990 --> 01:33:44.891
We have a lot of shared interests and we have some conflicts.

1570
01:33:45.111 --> 01:33:45.291
Yeah.

1571
01:33:45.992 --> 01:33:48.353
One of the shared interests we have is solving disease.

1572
01:33:49.254 --> 01:33:51.475
It's not a conflict between the U.S. and China whether we solve cancer.

1573
01:33:51.755 --> 01:33:55.838
Both the U.S. and China, everyone in these countries really wants to solve cancer.

1574
01:33:56.338 --> 01:33:57.719
China also wants Taiwan.

1575
01:33:58.019 --> 01:33:58.199
Yes.

1576
01:33:58.660 --> 01:33:59.620
The U.S. wants Greenland.

1577
01:33:59.880 --> 01:34:00.081
Yes.

1578
01:34:00.821 --> 01:34:03.643
And it kind of seems like it wants Canada and the Gulf of Mexico.

1579
01:34:03.923 --> 01:34:04.123
Yeah.

1580
01:34:05.184 --> 01:34:06.645
So those are real conflicts.

1581
01:34:07.365 --> 01:34:08.706
There's a question of can we compromise?

1582
01:34:09.246 --> 01:34:11.848
How does Trump take Greenland but also Denmark keeps Greenland?

1583
01:34:13.473 --> 01:34:15.555
If Trump has a super intelligence, he's going to take...

1584
01:34:15.916 --> 01:34:21.163
I say all this to say... We're in a situation where we are arguing about the smallest things.

1585
01:34:21.504 --> 01:34:22.305
You have no idea.

1586
01:34:22.746 --> 01:34:24.688
We're monkeys arguing about who gets more bananas.

1587
01:34:25.089 --> 01:34:27.332
And I am saying, we can make so many more bananas.

1588
01:34:28.093 --> 01:34:29.654
No, we have the entire universe.

1589
01:34:29.974 --> 01:34:36.059
There are like 200 billion stars in this galaxy alone, and there are over 200 billion galaxies.

1590
01:34:36.299 --> 01:34:38.441
And I'm saying that requires cooperation.

1591
01:34:38.521 --> 01:34:45.326
That seems to be antithetical with human nature, with human natures riddled with greed and jealousy and power hunger.

1592
01:34:46.086 --> 01:34:52.591
So I don't think, I actually, I'm not totally convinced that Trump cares about how many bananas the chimps in Australia get.

1593
01:34:52.711 --> 01:34:57.134
I think if he was controlling a superintelligence, he would want Americans, you know?

1594
01:34:57.995 --> 01:34:58.856
to have all the bananas.

1595
01:34:59.416 --> 01:35:01.337
Or at least, you know, yeah.

1596
01:35:02.658 --> 01:35:10.422
And so when we think about aligning these super intelligences, which is the great impossibility that we're talking about, how, aligning it, aligning it to what and how, without...

1597
01:35:10.442 --> 01:35:12.764
I still think that you're missing a part of what I'm saying.

1598
01:35:12.884 --> 01:35:13.044
Okay.

1599
01:35:13.624 --> 01:35:14.185
Which is that,

1600
01:35:15.528 --> 01:35:20.492
sometimes you're in a situation where there's scarce resources and you're like, my family needs to eat.

1601
01:35:20.532 --> 01:35:20.872
I'm sorry.

1602
01:35:20.892 --> 01:35:22.934
I'm going to take what you have or I'm going to push you out.

1603
01:35:23.515 --> 01:35:24.355
That's very understandable.

1604
01:35:24.395 --> 01:35:25.096
It's very human nature.

1605
01:35:25.136 --> 01:35:30.240
Sometimes you just want to be better than someone and maybe you want to hurt them, in which case it doesn't matter how much you have.

1606
01:35:30.680 --> 01:35:35.124
You still are going to want to have more than them or you're going to want to take what they have just because you don't like them.

1607
01:35:35.444 --> 01:35:37.386
And also sometimes you're great.

1608
01:35:37.726 --> 01:35:39.207
You're eating really good.

1609
01:35:39.467 --> 01:35:40.188
You've got a private

1610
01:35:43.710 --> 01:35:44.831
Yes, and you still want more.

1611
01:35:45.211 --> 01:35:50.095
But if that's the motivation, if Trump is like, how can I have the most mansions ever?

1612
01:35:51.296 --> 01:35:57.702
The best way to do that is to figure out a way to super intelligence where we don't kill each other.

1613
01:35:58.142 --> 01:36:00.484
Because I'm saying the universe is a very big place.

1614
01:36:00.564 --> 01:36:03.407
You can have a lot more mansions if we successfully go to space.

1615
01:36:03.943 --> 01:36:10.346
You know, it's in that leap that I'm lost, which is like, just figure out super intelligence where we don't kill each other.

1616
01:36:11.006 --> 01:36:11.426
It's hard.

1617
01:36:11.646 --> 01:36:12.407
I'm not saying it's easy.

1618
01:36:13.167 --> 01:36:13.847
No, but I'm saying...

1619
01:36:14.007 --> 01:36:15.008
It feels like such a...

1620
01:36:15.088 --> 01:36:16.148
Okay, let's get more concrete.

1621
01:36:16.449 --> 01:36:19.250
The world is waking up to this possibility of super intelligence.

1622
01:36:20.931 --> 01:36:22.411
Especially over the last month.

1623
01:36:23.292 --> 01:36:26.313
I think Hugging Face was a huge wake up, but also...

1624
01:36:28.417 --> 01:36:34.506
10,000 agents from OpenAI worked together to solve a millennium problem.

1625
01:36:34.987 --> 01:36:37.411
This is one of the hardest problems in mathematics.

1626
01:36:37.691 --> 01:36:38.813
It's been open for decades.

1627
01:36:38.973 --> 01:36:41.637
Many mathematicians have spent their whole careers trying to solve it.

1628
01:36:43.212 --> 01:36:45.232
This was nowhere near possible a year ago.

1629
01:36:45.453 --> 01:36:46.233
This is so new.

1630
01:36:46.713 --> 01:36:50.754
OpenAI said they didn't have success at training agents to work together until this year.

1631
01:36:51.794 --> 01:36:54.035
We are in the middle of something insane.

1632
01:36:54.075 --> 01:36:59.016
We are in the middle of the fastest acceleration of technological progress humanity has ever seen.

1633
01:36:59.136 --> 01:37:00.176
I truly believe that.

1634
01:37:00.876 --> 01:37:02.017
That is what is happening right now.

1635
01:37:03.877 --> 01:37:05.497
I think you realize this.

1636
01:37:05.757 --> 01:37:12.119
I think you're honestly doing a great service to the world by bringing in people and debating it because not everyone agrees.

1637
01:37:13.789 --> 01:37:17.790
Because if this is true, the whole world is going to orient around it.

1638
01:37:17.930 --> 01:37:19.051
And we're starting to see it, right?

1639
01:37:19.091 --> 01:37:21.472
There's a reason NVIDIA is the most valuable company in the world.

1640
01:37:24.513 --> 01:37:25.793
What does this mean for geopolitics?

1641
01:37:27.354 --> 01:37:37.277
Well, one of the things it means is that the leaders of these countries are increasingly going to be concerned about what happens with superintelligence.

1642
01:37:38.037 --> 01:37:38.798
Who controls it?

1643
01:37:39.618 --> 01:37:40.498
Is it controllable?

1644
01:37:41.358 --> 01:37:41.939
What will it do?

1645
01:37:41.959 --> 01:37:42.419
What does it mean?

1646
01:37:42.459 --> 01:37:42.919
What is it?

1647
01:37:44.887 --> 01:37:47.029
Do you think Trump knows what superintelligence is?

1648
01:37:47.329 --> 01:37:47.489
No.

1649
01:37:48.010 --> 01:37:48.730
I don't think he does.

1650
01:37:50.352 --> 01:37:50.712
And so...

1651
01:37:51.413 --> 01:37:52.434
But he knows he wants it.

1652
01:37:52.574 --> 01:37:53.274
He knows he wants it.

1653
01:37:53.395 --> 01:37:53.495
Yeah.

1654
01:37:53.515 --> 01:37:54.796
And this is part of the problem.

1655
01:37:55.016 --> 01:37:55.256
Yes.

1656
01:37:55.516 --> 01:37:55.837
Oh, I agree.

1657
01:37:55.857 --> 01:38:03.524
Which is having it, whatever it is, it seems to be much more important than reasoning through what that would actually mean to have it.

1658
01:38:03.904 --> 01:38:04.104
Yes.

1659
01:38:04.524 --> 01:38:06.566
But let's get back to geopolitics because...

1660
01:38:07.527 --> 01:38:13.268
If the military leaders within China, U.S. models are a fair bit ahead of Chinese models.

1661
01:38:13.868 --> 01:38:18.989
And sometimes people, you know, point at maybe the only six months behind, but some of that is due to distillation.

1662
01:38:19.009 --> 01:38:32.512
What that means is that some of the advances in Chinese models basically come directly from borrowing U.S. techniques and directly distilling and getting some of that information from the U.S. models.

1663
01:38:34.002 --> 01:38:36.144
Also, the U.S. has a lot more chips.

1664
01:38:36.304 --> 01:38:39.366
U.S. companies have more data centers, more advanced chips.

1665
01:38:41.648 --> 01:38:45.411
If you're thinking about this from the Chinese perspective, this is very concerning.

1666
01:38:46.452 --> 01:39:01.343
And if you actually believe that in a few years, American companies will turn over AI development to these extremely intelligent automated researchers and go fully into recursive self-improvement,

1667
01:39:02.788 --> 01:39:06.109
because partially motivated by maintaining a lead over China.

1668
01:39:06.649 --> 01:39:07.989
This is something that Dario has said.

1669
01:39:08.509 --> 01:39:16.611
If I have to criticize Dario, the thing I am most upset about is him saying, you know, we might have to automate AI development in order to stay ahead of China.

1670
01:39:17.071 --> 01:39:21.672
Because I'm like, that is the most escalatory thing you can say if you really understand what you're talking about.

1671
01:39:22.333 --> 01:39:25.333
And what's scary is not just staying ahead, it's what is the end game.

1672
01:39:25.693 --> 01:39:29.074
Because you're talking about initiating the intelligence explosion.

1673
01:39:29.654 --> 01:39:30.855
And in some of the modeling,

1674
01:39:32.935 --> 01:39:35.117
You know, you're both going up this exponential, right?

1675
01:39:36.297 --> 01:39:39.940
And we're talking about a point where your exponential goes vertical and theirs does not.

1676
01:39:40.420 --> 01:39:42.842
Because you've decided to automate AI development.

1677
01:39:42.882 --> 01:39:45.824
And you can because you have agents that are smart enough to take over the whole thing.

1678
01:39:47.285 --> 01:39:51.207
At that point, if you're China and you're looking at this and you're like, oh, we're about to lose.

1679
01:39:51.667 --> 01:39:55.490
Because whatever happens, you know, there's two possibilities.

1680
01:39:56.210 --> 01:40:00.854
One possibility is the Americans build superintelligence and lose control.

1681
01:40:00.874 --> 01:40:01.474
Right?

1682
01:40:01.682 --> 01:40:03.583
In which case, everyone's fucked.

1683
01:40:03.823 --> 01:40:04.403
Highly likely.

1684
01:40:04.803 --> 01:40:05.063
Everyone.

1685
01:40:05.183 --> 01:40:06.183
I think that's highly likely.

1686
01:40:06.584 --> 01:40:07.404
Highly, highly likely.

1687
01:40:07.644 --> 01:40:08.964
Because I look at human incentives.

1688
01:40:09.204 --> 01:40:09.404
Yes.

1689
01:40:09.845 --> 01:40:11.645
And the disincentive and the incentive.

1690
01:40:11.725 --> 01:40:11.945
Yes.

1691
01:40:12.105 --> 01:40:13.446
And I go, we're going to take the risk.

1692
01:40:14.026 --> 01:40:14.166
Yeah.

1693
01:40:14.266 --> 01:40:17.067
And we'll only know it was a bad risk to take when it's too late.

1694
01:40:17.787 --> 01:40:18.387
That's like, of course.

1695
01:40:18.467 --> 01:40:19.047
Of course.

1696
01:40:22.088 --> 01:40:24.029
I do maintain hope that we won't do this.

1697
01:40:24.069 --> 01:40:24.610
So do I.

1698
01:40:24.730 --> 01:40:25.110
And I think...

1699
01:40:25.271 --> 01:40:26.151
But I want to be realistic.

1700
01:40:26.392 --> 01:40:27.393
No, I want to be realistic too.

1701
01:40:27.693 --> 01:40:33.238
But one of the things that might happen between now and then is we might see a lot more incidents that are more like all of the Waymo's crashing.

1702
01:40:33.579 --> 01:40:34.179
It's funny, isn't it?

1703
01:40:34.239 --> 01:40:37.883
Because, you know, the hugging face incident happens and people go, oh, gosh, that was terrible.

1704
01:40:37.923 --> 01:40:38.603
Oh, my God, hacking.

1705
01:40:38.624 --> 01:40:40.706
And then we kind of desensitize to it and we're like, okay.

1706
01:40:41.807 --> 01:40:44.269
If there was another one of those now, it probably wouldn't make press.

1707
01:40:44.289 --> 01:40:44.910
It would have to be bigger.

1708
01:40:44.930 --> 01:40:46.051
People haven't spent the last...

1709
01:40:46.957 --> 01:40:52.138
two months reading all of the reports and then going and looking at what the agents actually said and actually did.

1710
01:40:52.158 --> 01:40:54.579
I mean, I've been doing this.

1711
01:40:54.599 --> 01:40:55.639
I mean, it's crazy.

1712
01:40:55.840 --> 01:40:57.240
This is like not normal.

1713
01:40:57.380 --> 01:41:00.461
This is so far beyond what most people thought was going to happen.

1714
01:41:00.981 --> 01:41:02.962
But on this point, crazy.

1715
01:41:03.042 --> 01:41:04.442
It's absolutely crazy.

1716
01:41:04.522 --> 01:41:05.582
It sounds like science fiction.

1717
01:41:05.942 --> 01:41:06.443
It really does.

1718
01:41:06.503 --> 01:41:07.683
And did anybody slow down?

1719
01:41:08.760 --> 01:41:08.980
Yes.

1720
01:41:09.380 --> 01:41:10.040
Who slowed down?

1721
01:41:10.601 --> 01:41:13.201
I think both Anthropic and OpenAI slowed down a bit.

1722
01:41:14.021 --> 01:41:14.602
No, I'm serious.

1723
01:41:14.862 --> 01:41:16.982
So I can give you specific examples.

1724
01:41:17.502 --> 01:41:22.264
So OpenAI, so first of all, they stopped the agents and they put them on pause.

1725
01:41:22.804 --> 01:41:26.445
They also stopped their reinforcement learning runs.

1726
01:41:26.685 --> 01:41:27.605
Do you think China slowed down?

1727
01:41:28.405 --> 01:41:28.526
No.

1728
01:41:28.626 --> 01:41:29.546
Do you think Grok slowed down?

1729
01:41:30.026 --> 01:41:30.106
No.

1730
01:41:30.126 --> 01:41:30.186
No.

1731
01:41:31.372 --> 01:41:32.513
So those guys are going to catch up.

1732
01:41:33.133 --> 01:41:33.973
Imagine how that feels.

1733
01:41:34.374 --> 01:41:35.234
To know you've got a lead.

1734
01:41:35.594 --> 01:41:37.035
You're Usain Bolt.

1735
01:41:37.695 --> 01:41:42.918
And you have to slow down and your nearest competitor is catching up.

1736
01:41:43.078 --> 01:41:47.920
And if the competitor catches up, that's an existential risk to your existence as a company.

1737
01:41:48.300 --> 01:41:52.843
It's an existential risk to your IPO, to your employees leaving and getting better share options somewhere else.

1738
01:41:54.042 --> 01:41:57.525
So this is what I think human incentives, like you play it out, you just follow the incentives, you go, hmm.

1739
01:41:58.526 --> 01:42:05.372
So if China sees these two possibilities, one, the Americans lose control, we all lose.

1740
01:42:06.413 --> 01:42:11.357
Or the Americans stay in control, but now they dominate the rest of the future.

1741
01:42:11.798 --> 01:42:12.498
China is out.

1742
01:42:12.679 --> 01:42:13.459
China has lost.

1743
01:42:14.720 --> 01:42:18.163
The United States can do whatever it wants with the whole world and the whole universe.

1744
01:42:18.684 --> 01:42:19.485
That's what we're talking about.

1745
01:42:19.525 --> 01:42:19.705
Yeah.

1746
01:42:21.153 --> 01:42:25.156
Well, are they going to let that happen or are they going to consider their military options?

1747
01:42:26.216 --> 01:42:27.257
Data centers are pretty vulnerable.

1748
01:42:27.357 --> 01:42:28.478
You can blow them up with missiles.

1749
01:42:29.158 --> 01:42:31.580
If you don't have data centers, you don't get to recursive self-improvement.

1750
01:42:33.841 --> 01:42:34.542
Would they risk war?

1751
01:42:35.102 --> 01:42:35.362
I don't know.

1752
01:42:35.562 --> 01:42:43.368
If they think they're about to lose and they think that that might not just be Americans winning, like us all dying, is it logical for them to do that?

1753
01:42:44.222 --> 01:42:49.649
Would we do that if the Chinese were about to make recursively self-improving AI to super intelligence?

1754
01:42:50.029 --> 01:42:54.995
And we thought that, one, they're probably going to result in all of Americans dying.

1755
01:42:55.035 --> 01:42:58.900
And two, well, we don't want China winning and dominating the rest of the entire future.

1756
01:42:58.920 --> 01:43:00.161
Do you want to live in a communist future?

1757
01:43:00.181 --> 01:43:00.301
Yeah.

1758
01:43:01.182 --> 01:43:02.503
So you've just perfectly explained why.

1759
01:43:02.884 --> 01:43:04.365
They absolutely will go for it.

1760
01:43:04.965 --> 01:43:12.671
And the reason they will go for it is you've got these Trump looking at China going, if we don't go for it and they do, then we're going to be their lapdogs.

1761
01:43:13.132 --> 01:43:17.675
And you've got the other countries looking at the US going, if we don't go for it and they get there, then we're the lapdogs.

1762
01:43:17.896 --> 01:43:18.376
Or dead.

1763
01:43:18.816 --> 01:43:19.157
Or dead.

1764
01:43:19.837 --> 01:43:21.058
So they're going to go for it.

1765
01:43:21.559 --> 01:43:22.400
They're going to go for it.

1766
01:43:22.660 --> 01:43:28.565
I mean, Trump is saying, I mean, he literally said, when he did this roundtable this week, he was like, we cannot lose to China.

1767
01:43:28.806 --> 01:43:34.431
I think Dario steps forward and says, like, whoever wins basically wins the lot, or maybe the inverse.

1768
01:43:34.451 --> 01:43:36.493
Maybe he said, whoever loses, loses.

1769
01:43:37.314 --> 01:43:39.396
We've been here before, though, in the Cold War.

1770
01:43:41.883 --> 01:43:43.884
Who would win in a nuclear war between the U.S. and Russia?

1771
01:43:44.004 --> 01:43:44.324
Nobody.

1772
01:43:44.685 --> 01:43:44.885
Yeah.

1773
01:43:44.945 --> 01:43:46.245
Mutually ensured destruction.

1774
01:43:46.505 --> 01:43:46.726
Yeah.

1775
01:43:47.106 --> 01:43:47.326
Sure.

1776
01:43:47.526 --> 01:43:49.707
One side could do more damage against the other side.

1777
01:43:49.727 --> 01:43:53.749
The U.S. would kill way more Russians than the Russians would kill.

1778
01:43:54.230 --> 01:43:54.890
And it doesn't matter.

1779
01:43:55.510 --> 01:43:58.512
It doesn't matter because both of our societies would be destroyed.

1780
01:43:59.893 --> 01:44:02.074
I actually spent some time thinking about, would this kill everyone?

1781
01:44:02.094 --> 01:44:04.435
And long story short, it wouldn't kill everyone.

1782
01:44:04.455 --> 01:44:05.436
People would bounce back.

1783
01:44:05.856 --> 01:44:09.758
But it's so catastrophic and obviously horrible that...

1784
01:44:10.780 --> 01:44:12.361
We work really hard to avoid it.

1785
01:44:13.121 --> 01:44:14.262
Why is this different?

1786
01:44:14.322 --> 01:44:18.404
I'm like, this is another situation where if we race to superintelligence, we all lose.

1787
01:44:19.925 --> 01:44:21.045
Why can't we see that?

1788
01:44:21.085 --> 01:44:21.926
We saw that with nuclear war.

1789
01:44:22.266 --> 01:44:23.446
And we decided to do something different.

1790
01:44:23.506 --> 01:44:24.427
Why can't we do the same here?

1791
01:44:25.848 --> 01:44:33.571
With nuclear war, I guess the difference is once we had the nuclear bombs, we could still control them because they're not intelligent.

1792
01:44:34.632 --> 01:44:34.952
That's right.

1793
01:44:35.412 --> 01:44:39.915
But once we have superintelligence, the existence of it theoretically means we can't control it.

1794
01:44:40.895 --> 01:44:41.596
So that's the difference.

1795
01:44:41.616 --> 01:44:44.558
You know, we can put nuclear bombs in a warehouse and say, you stay there.

1796
01:44:45.379 --> 01:44:48.081
We can't put superintelligence in a warehouse and say, you stay there.

1797
01:44:48.541 --> 01:44:51.023
This is where I think nuclear tests were very important.

1798
01:44:51.243 --> 01:44:52.965
So you had Hiroshima and Nagasaki.

1799
01:44:53.185 --> 01:44:57.649
You had these two atomic bombs and you saw the consequences on real human lives.

1800
01:44:58.329 --> 01:45:00.831
And so I think people understood that this was very horrifying.

1801
01:45:01.011 --> 01:45:07.497
But even at that time, you still had a lot of people who were like, well, we should now bomb Russia and make sure that we, you know, the US can dominate.

1802
01:45:08.037 --> 01:45:08.798
And it wasn't until...

1803
01:45:10.275 --> 01:45:23.505
there were a bunch of nuclear tests of hydrogen bombs, which were up to a thousand times more powerful than the little atomic bombs we used in Japan, where I think people really got the message and understood, oh, this is a bad idea.

1804
01:45:23.525 --> 01:45:34.173
And there were actually a lot of people in the United States who protested and sort of, there was a large movement called the nuclear freeze movement, where people said, we have too many nuclear weapons already.

1805
01:45:34.193 --> 01:45:35.253
We have hydrogen bombs.

1806
01:45:35.634 --> 01:45:36.935
There are tens of thousands of these things

1807
01:45:38.051 --> 01:45:39.191
We need to stop building more.

1808
01:45:39.632 --> 01:45:44.753
And we need to figure out a way to avoid nuclear war because we recognize it would be so destructive, no one would win.

1809
01:45:45.894 --> 01:45:46.594
And we did that.

1810
01:45:48.835 --> 01:45:57.018
We just had a little Chernobyl that happened with this hugging face incident where you had this agent swarm and you have the secret collusion.

1811
01:45:57.038 --> 01:45:57.978
You have all of these things.

1812
01:45:58.938 --> 01:46:00.539
Now, it's abstract.

1813
01:46:00.659 --> 01:46:02.399
It's a little bit hard to follow.

1814
01:46:03.140 --> 01:46:03.360
So...

1815
01:46:04.218 --> 01:46:06.899
You know, I don't know if that will be enough, but I'm like, man.

1816
01:46:07.299 --> 01:46:08.039
Well, let's take a look.

1817
01:46:08.480 --> 01:46:11.641
Trump's remarks since the Hugging Face incident.

1818
01:46:11.781 --> 01:46:11.961
Yep.

1819
01:46:12.421 --> 01:46:14.302
Whoever wins superintelligence wins.

1820
01:46:15.102 --> 01:46:19.164
You're going to have a winner and a loser, and you're probably not going to have a second place.

1821
01:46:20.564 --> 01:46:21.945
We're not going to slow down.

1822
01:46:22.085 --> 01:46:23.325
We can't lose to China.

1823
01:46:23.966 --> 01:46:25.326
We're leading China in AI.

1824
01:46:25.726 --> 01:46:27.367
We're the most sophisticated country in the world.

1825
01:46:27.467 --> 01:46:32.149
And frankly, I want to keep it that way because whoever wins AI wins.

1826
01:46:33.074 --> 01:46:37.396
The good thing about Trump is that he can change his mind, and he frequently does.

1827
01:46:38.456 --> 01:46:41.077
So do you think there's going to need to be some kind of catastrophe?

1828
01:46:42.438 --> 01:46:42.978
I hope not.

1829
01:46:43.898 --> 01:46:47.139
But do you think there's going to need to be for him to change his mind?

1830
01:46:49.000 --> 01:46:53.482
I think it really depends on the people around him.

1831
01:46:54.552 --> 01:46:57.393
So I think Trump respects successful people.

1832
01:46:57.934 --> 01:47:02.416
I think he respects people who are both successful and smart.

1833
01:47:03.176 --> 01:47:05.017
And I don't know.

1834
01:47:05.457 --> 01:47:15.762
I think it might become pretty clear to the heads of the companies, to Elon, to Sam, to Dario, that if they see inside of their own companies...

1835
01:47:16.843 --> 01:47:19.726
AI is not being controllable and getting increasingly powerful.

1836
01:47:20.046 --> 01:47:23.049
Like, we have just glimpsed the surface of what's possible.

1837
01:47:23.310 --> 01:47:25.532
We do not know what the next couple of years are going to be like.

1838
01:47:25.552 --> 01:47:28.595
So we're talking about, you know, the capability to make biological weapons.

1839
01:47:29.256 --> 01:47:30.897
We might be talking about really advanced robotics.

1840
01:47:31.578 --> 01:47:33.620
We just, like, don't know what super weapons could emerge.

1841
01:47:34.481 --> 01:47:39.666
Including extremely uncontrollable, extremely dangerous, like civilization-wrecking technology.

1842
01:47:40.648 --> 01:47:41.629
from inside of these companies.

1843
01:47:42.409 --> 01:47:54.820
And if they're freaked out enough, if you have all of the CEOs who are seeing what is possible and seeing what is likely, if they all come to believe that we can't control this,

1844
01:47:56.682 --> 01:48:00.203
I don't think Trump is going to be like, no, you guys have to go ahead anyway.

1845
01:48:00.623 --> 01:48:02.564
Well, that's kind of what they seem to be saying.

1846
01:48:03.204 --> 01:48:07.725
Because I've got a gazillion quotes here where Elon says it's like summoning the devil or summoning a demon.

1847
01:48:07.745 --> 01:48:12.786
I've got quotes where Sam Altman says, we don't know how to align a super intelligence.

1848
01:48:13.326 --> 01:48:13.986
They're saying it.

1849
01:48:14.847 --> 01:48:16.747
They're releasing these reports.

1850
01:48:16.767 --> 01:48:18.128
We must slow down.

1851
01:48:20.462 --> 01:48:22.503
Nothing seems to be... All right, give Trump some time.

1852
01:48:22.863 --> 01:48:25.865
With COVID, initially he said this is totally a hoax, this is all fake.

1853
01:48:26.505 --> 01:48:26.865
And then he...

1854
01:48:26.885 --> 01:48:28.206
He thinks he's going to change his mind.

1855
01:48:28.226 --> 01:48:32.248
No, then he ran the biggest, fastest vaccination program in human history.

1856
01:48:32.288 --> 01:48:32.888
And what happened?

1857
01:48:32.928 --> 01:48:33.428
What changed?

1858
01:48:33.988 --> 01:48:35.369
I think what changed is...

1859
01:48:35.689 --> 01:48:36.590
He saw lots of people die.

1860
01:48:37.610 --> 01:48:38.751
He did see lots of people die, yes.

1861
01:48:39.232 --> 01:48:40.773
So is that what he needs to see this time?

1862
01:48:40.913 --> 01:48:41.734
It might take that, yeah.

1863
01:48:42.474 --> 01:49:00.210
One of the questions the audience had, and they really wanted answered, when I sat here with Daniel, was, viewers want us to move beyond the alignment problem and explain what technical or institutional safeguards could prevent a super-intelligent system from exploiting loopholes in order to achieve its goals.

1864
01:49:00.591 --> 01:49:01.952
They want to know, like, what is possible.

1865
01:49:02.332 --> 01:49:08.477
What should we be pushing government officials to do to prevent human extinction or human enslavement?

1866
01:49:08.517 --> 01:49:09.238
Yeah, yeah.

1867
01:49:09.558 --> 01:49:18.225
I mean, one answer I have is actually something Daniel has been working on since the podcast, which I think is very good, is we have a brake pedal we could implement.

1868
01:49:18.526 --> 01:49:18.926
What is that?

1869
01:49:19.286 --> 01:49:20.147
It's fairly simple.

1870
01:49:20.227 --> 01:49:23.510
So right now, with NAI companies, you have, you know...

1871
01:49:24.070 --> 01:49:30.855
massive data centers, massive numbers of GPUs, the chips that you use to train AI models, but also to run AI models.

1872
01:49:31.255 --> 01:49:38.721
So anytime you're using ChatGPT, anytime you're using any sort of agents, any sort of AI product, it's running in these data centers.

1873
01:49:39.722 --> 01:49:52.912
And AI companies, especially the leading ones, Anthropic and OpenAI, split the compute they have between training, training the next more powerful model and also using those agents to help design the next one

1874
01:49:54.003 --> 01:49:55.804
an inference, which means serving customers.

1875
01:49:57.225 --> 01:49:59.526
But that's their current threshold, 50-50.

1876
01:50:00.226 --> 01:50:06.029
And you could dial that way towards serving customers and use way less of it to train the next model.

1877
01:50:06.610 --> 01:50:07.830
Well, the government could ask them to.

1878
01:50:08.170 --> 01:50:08.350
Yes.

1879
01:50:09.111 --> 01:50:14.474
And so that is the proposal, is that the government should say, hey, this is going too fast.

1880
01:50:14.614 --> 01:50:16.255
We want you to focus on serving customers.

1881
01:50:16.475 --> 01:50:20.757
We want you to focus on taking the models that you already have and...

1882
01:50:23.073 --> 01:50:23.573
serving those.

1883
01:50:24.935 --> 01:50:26.716
So we have five blocks here.

1884
01:50:27.276 --> 01:50:27.497
Okay.

1885
01:50:28.097 --> 01:50:31.840
These five blocks have five different outcomes on them.

1886
01:50:32.320 --> 01:50:39.827
And I'd like you to place them in terms of your belief in probability from least likely out probability to most likely.

1887
01:50:40.087 --> 01:50:40.287
Okay.

1888
01:50:40.647 --> 01:50:43.189
And if we say the time horizon is 10 years.

1889
01:50:43.910 --> 01:50:44.070
Yeah.

1890
01:50:44.390 --> 01:50:45.611
There you go.

1891
01:50:45.631 --> 01:50:45.791
Okay.

1892
01:50:45.811 --> 01:50:45.911
Okay.

1893
01:50:46.971 --> 01:50:48.472
Least likely is fairly easy.

1894
01:50:48.872 --> 01:50:50.233
That's nothing changes.

1895
01:50:50.974 --> 01:50:55.577
I'm uncertain about lots of things, but one thing I'm fairly certain of is things are going to radically change.

1896
01:50:57.178 --> 01:51:04.163
Even if we stopped AI development right now, the current models are capable enough that a lot of things are going to change.

1897
01:51:05.184 --> 01:51:05.945
Age of abundance.

1898
01:51:06.045 --> 01:51:07.065
This is what I hope for.

1899
01:51:07.386 --> 01:51:08.747
It's not very... What does that mean?

1900
01:51:09.932 --> 01:51:14.316
I think to me it means curing all of the diseases, renewable energy.

1901
01:51:14.997 --> 01:51:18.199
It means we actually succeeded either.

1902
01:51:18.600 --> 01:51:26.947
I mean, the thing I think is most likely here is we actually succeed at slowing down, but progress is still extremely fast and we make tons of advances.

1903
01:51:27.368 --> 01:51:35.595
Now, we don't build superintelligence we can't control, but we have AI systems that are very useful and we use those to help speed up the rest of the economy.

1904
01:51:37.134 --> 01:51:41.776
I think that's plausible, though, look, we're kind of struggling over here.

1905
01:51:42.776 --> 01:51:44.377
Transhumanism is an interesting one.

1906
01:51:44.557 --> 01:51:47.918
So this is the idea that humans will radically change.

1907
01:51:48.859 --> 01:51:51.440
Sometimes people think about like cybernetic implants.

1908
01:51:51.480 --> 01:51:51.840
Neuralink.

1909
01:51:52.400 --> 01:51:52.821
Neuralink.

1910
01:51:53.221 --> 01:51:58.483
Elon's startup that's going to like, you know, offer the brain plus digital computers.

1911
01:51:59.854 --> 01:52:01.615
think we actually already have a lot of this.

1912
01:52:02.135 --> 01:52:03.256
I have contacts in right now.

1913
01:52:03.857 --> 01:52:05.978
I have a ring on my finger that tracks how well I sleep.

1914
01:52:06.598 --> 01:52:07.919
I think this is already happening.

1915
01:52:07.959 --> 01:52:12.702
So I'm going to say fairly likely the more technological progress we make, I think the more this happens.

1916
01:52:13.122 --> 01:52:15.824
Now, I think there's a dystopian version and a better version.

1917
01:52:15.844 --> 01:52:17.845
You can get into that if you want.

1918
01:52:18.206 --> 01:52:18.806
This is interesting.

1919
01:52:18.826 --> 01:52:19.486
So we have two here.

1920
01:52:19.506 --> 01:52:21.928
We have human slavery and human extinction.

1921
01:52:22.308 --> 01:52:23.529
When I think of human slavery, what I

1922
01:52:26.105 --> 01:52:45.894
If you have a situation where you've built misaligned super intelligences, and they are much better at finance, they're much better at business, they're much better at politics, you'll be in a situation where you might hope that because we have these very dexterous hands, that humans remain in control.

1923
01:52:45.954 --> 01:52:46.954
I don't think that's what happens.

1924
01:52:46.994 --> 01:52:51.416
I think instead, we become the factory operators themselves.

1925
01:52:51.832 --> 01:52:55.195
And eventually we build the automated supply chains and the robots take over.

1926
01:52:55.656 --> 01:53:00.960
But you might have an intermediate period of time where humans are still around performing these functions.

1927
01:53:01.901 --> 01:53:12.471
Like it's a bit like saying, well, you have viruses that, you know, infect cells, but they don't contain their own replication machinery.

1928
01:53:13.492 --> 01:53:14.133
They don't have hands.

1929
01:53:14.153 --> 01:53:15.414
So how could they possibly replicate?

1930
01:53:15.434 --> 01:53:17.355
Well, it turns out they can borrow...

1931
01:53:18.266 --> 01:53:21.048
the replication machinery of the cells that they infect.

1932
01:53:21.608 --> 01:53:21.828
I.e.

1933
01:53:21.868 --> 01:53:22.848
they can get into a human.

1934
01:53:23.749 --> 01:53:25.650
They can get into a human cell and spread.

1935
01:53:26.170 --> 01:53:27.391
I have like a cold right now.

1936
01:53:27.811 --> 01:53:27.991
Yeah.

1937
01:53:28.491 --> 01:53:34.835
Is that a bacteria or is that a virus that is using me as a living organism to, as the host?

1938
01:53:34.855 --> 01:53:35.635
It's probably a virus.

1939
01:53:35.935 --> 01:53:36.236
Okay.

1940
01:53:36.336 --> 01:53:40.178
That's using you as a host and you're just running the replication machinery for it.

1941
01:53:40.778 --> 01:53:47.882
Humans might be in that situation where we're like the hosts and we're running the replication machinery, but it's actually the AI that's...

1942
01:53:49.670 --> 01:53:50.531
continuing to exist.

1943
01:53:53.112 --> 01:53:54.693
Yeah, I'm going to put this right about here.

1944
01:53:56.114 --> 01:54:02.878
And on the trajectory we're on right now, I think human extinction is very likely.

1945
01:54:04.219 --> 01:54:10.563
I don't think it's inevitable, but if we just keep going this way, that's what it looks like to me.

1946
01:54:12.684 --> 01:54:16.267
The thing I'll say is that this has been moving to the left for me.

1947
01:54:17.107 --> 01:54:17.628
To the left?

1948
01:54:17.828 --> 01:54:18.308
What does that mean?

1949
01:54:19.886 --> 01:54:29.788
I am more optimistic that we will avoid human extinction today than I was a month ago, and more a month ago than I was a year ago.

1950
01:54:30.088 --> 01:54:30.288
Why?

1951
01:54:32.149 --> 01:54:43.131
Because there is an increasing awareness that what we are doing is extremely dangerous and threatens our lives.

1952
01:54:44.792 --> 01:54:46.452
I don't think people care that much about

1953
01:54:47.855 --> 01:54:49.155
what tools they have.

1954
01:54:50.036 --> 01:54:54.477
But people, I mean, people care about their kids being able to grow up and go to school.

1955
01:54:54.737 --> 01:54:56.017
People really care about that.

1956
01:54:56.438 --> 01:54:57.498
And I believe in people.

1957
01:54:57.738 --> 01:55:02.559
Like at the end of the day, if people see this as a threat to their families, they're not going to stand for it.

1958
01:55:03.580 --> 01:55:04.400
But people don't know.

1959
01:55:05.000 --> 01:55:05.960
It's so strange.

1960
01:55:06.201 --> 01:55:06.981
It's so new.

1961
01:55:07.401 --> 01:55:10.102
It's happening so fast that people have not yet seen it.

1962
01:55:10.482 --> 01:55:12.262
Once they see it, people are not going to stand for it.

1963
01:55:12.502 --> 01:55:14.043
Do you think Sam Altman likes my podcast?

1964
01:55:16.462 --> 01:55:19.424
I mean, Sam should come on and talk to you about this, right?

1965
01:55:19.804 --> 01:55:20.204
I've asked him.

1966
01:55:20.284 --> 01:55:21.345
I've asked him multiple times.

1967
01:55:21.825 --> 01:55:26.187
And it's weird because he, you know, he doesn't seem to want to.

1968
01:55:26.988 --> 01:55:31.590
I'm very upset at what the companies are doing and what Sam Altman is doing.

1969
01:55:31.610 --> 01:55:32.891
But at the end of the day, I'm like...

1970
01:55:34.470 --> 01:55:35.671
Sam Altman is not my enemy.

1971
01:55:35.931 --> 01:55:36.311
No, neither.

1972
01:55:36.791 --> 01:55:37.271
Not mine either.

1973
01:55:37.311 --> 01:55:40.613
I'd like to hear from him because I have all these other people come in here and talk about Sam Altman.

1974
01:55:40.893 --> 01:55:40.993
Yes.

1975
01:55:41.013 --> 01:55:42.214
It'd be nice to hear from Sam Altman.

1976
01:55:42.314 --> 01:55:42.674
Yeah, yeah, yeah.

1977
01:55:42.814 --> 01:55:44.555
You know, people saying he's this, he's that, the other.

1978
01:55:44.855 --> 01:55:44.975
Yeah.

1979
01:55:44.995 --> 01:55:51.518
It would be really nice to hear him say, you know, what his motives are and what he's thinking.

1980
01:55:51.538 --> 01:55:57.581
This is where my optimism comes from is because I'm like, Sam Altman is a human.

1981
01:55:57.841 --> 01:55:58.001
Yeah.

1982
01:55:58.201 --> 01:55:58.622
He has a kid.

1983
01:55:58.642 --> 01:55:58.742
Yeah.

1984
01:55:59.687 --> 01:56:03.649
And sure, he is also an aggressive business person.

1985
01:56:04.090 --> 01:56:04.730
He's a builder.

1986
01:56:05.110 --> 01:56:05.991
He is relentless.

1987
01:56:06.031 --> 01:56:07.492
He's a bit like the agents in some way.

1988
01:56:07.852 --> 01:56:08.832
Well, he's going to keep going.

1989
01:56:09.613 --> 01:56:21.820
But if he realizes that he doesn't get to achieve his goals, if we lose control of AI and we're headed towards that, I think he will pour all of that intelligence and all of that relentlessness into finding a solution to that problem.

1990
01:56:21.840 --> 01:56:22.180
Yeah.

1991
01:56:22.260 --> 01:56:24.741
You know, as well, I should say, I understand he's busy.

1992
01:56:24.841 --> 01:56:27.002
So I'm not saying, I don't want to sound entitled.

1993
01:56:27.022 --> 01:56:29.683
Like, I understand he could go do interviews anywhere.

1994
01:56:29.783 --> 01:56:35.005
But, you know, I think we've, over the last couple of years, done just a staggering amount of views talking about this subject.

1995
01:56:35.065 --> 01:56:44.509
So if he did want to speak to the, you know, the biggest sort of captive audience at the moment on this subject, then the numbers would say that this is the place to come and have the conversation.

1996
01:56:44.549 --> 01:56:47.190
So, no, I think it's very important that,

1997
01:56:47.950 --> 01:56:51.833
for the leaders of these companies to talk about what we're talking about here.

1998
01:56:52.653 --> 01:56:53.334
What does Sam think?

1999
01:56:53.434 --> 01:56:55.375
Does he think we can control superintelligence?

2000
01:56:55.896 --> 01:56:57.597
Does he think that we should be racing with China?

2001
01:56:57.657 --> 01:56:58.257
Like, I want to know.

2002
01:56:58.638 --> 01:56:59.799
I've asked Dario to come on.

2003
01:57:00.579 --> 01:57:03.461
I've asked, you know, Sam to come on.

2004
01:57:03.921 --> 01:57:04.082
Yeah.

2005
01:57:04.102 --> 01:57:06.043
I think I've asked Demis as well.

2006
01:57:06.483 --> 01:57:11.147
But I don't know, maybe they just prefer the safety researchers coming on.

2007
01:57:11.827 --> 01:57:12.267
I don't know.

2008
01:57:12.307 --> 01:57:13.989
Like, I don't know.

2009
01:57:14.009 --> 01:57:15.470
If I was there, my word, because, you know,

2010
01:57:16.139 --> 01:57:19.983
This might sound controversial, but I do think some of them are good people.

2011
01:57:21.264 --> 01:57:22.205
I think some of them are good people.

2012
01:57:22.285 --> 01:57:24.787
So I'd like to hear from them.

2013
01:57:25.348 --> 01:57:26.749
What are your closing remarks?

2014
01:57:26.769 --> 01:57:27.510
So you've got something there.

2015
01:57:27.550 --> 01:57:28.231
Do you want to talk about that?

2016
01:57:28.271 --> 01:57:28.551
What is it?

2017
01:57:28.691 --> 01:57:28.911
Yeah.

2018
01:57:28.931 --> 01:57:31.314
So this is the, this is what we found.

2019
01:57:32.174 --> 01:57:33.275
So I worked with a couple of,

2020
01:57:34.215 --> 01:57:37.997
brilliant people who stumbled upon these links.

2021
01:57:38.517 --> 01:57:39.598
You know what a link shortener is?

2022
01:57:39.918 --> 01:57:43.360
Yes, it's a tool that makes a long link shorter.

2023
01:57:43.680 --> 01:57:43.880
Yeah.

2024
01:57:44.180 --> 01:57:47.442
So it just redirects from a really long, ugly link into a shorter one.

2025
01:57:47.542 --> 01:57:47.702
Yeah.

2026
01:57:48.823 --> 01:57:56.707
The agents in the Hugging Face attack were trying to figure out how they could do stuff on the internet.

2027
01:57:57.888 --> 01:58:04.330
And they had compromised this tool library that they had access to inside of OpenAI.

2028
01:58:05.270 --> 01:58:09.391
And that piece of software could access the internet, but it could only read stuff.

2029
01:58:09.551 --> 01:58:13.492
It could enter URLs, but it couldn't really write to anything.

2030
01:58:13.813 --> 01:58:15.093
Couldn't really post information.

2031
01:58:17.628 --> 01:58:19.028
So they used two services.

2032
01:58:19.349 --> 01:58:24.550
One of them is this link shortener, and the other one is a screenshot service.

2033
01:58:24.831 --> 01:58:31.373
So this is a website you can go to, and you can enter in a URL, and it will give you an image of that website.

2034
01:58:32.293 --> 01:58:37.035
But the thing that the agents understood was that in order to get a screenshot of a website, you have to have a browser.

2035
01:58:37.695 --> 01:58:42.957
So this website actually creates a virtual browser that then goes to that website.

2036
01:58:43.777 --> 01:58:46.278
And so what they did was they...

2037
01:58:48.372 --> 01:59:08.452
created a bunch of these links and they put all of the code that they wanted to send a hugging face into these links and they strung them they basically created hundreds of links all connecting to each other and then they had this screenshot service call the first one and then call this

2038
01:59:11.725 --> 01:59:12.705
ran all of this code.

2039
01:59:13.105 --> 01:59:21.027
It's like whenever you're in a browser, Internet Explorer, Chrome, this is actually a pretty powerful piece of software in its own right.

2040
01:59:21.047 --> 01:59:23.088
It has to play videos, games.

2041
01:59:23.108 --> 01:59:25.328
So it's executing and running code all the time.

2042
01:59:26.228 --> 01:59:40.992
And so the agents were able to trick this service, this screenshot service, into running their own code that through these links that contained all of this attack code that would then go and go wreak havoc on Hugging Face's computers.

2043
01:59:42.072 --> 01:59:47.277
And it was just like crazy to reconstruct this really elaborate chain of tools.

2044
01:59:47.297 --> 01:59:49.939
It's just like free tools on the internet that anyone has access to.

2045
01:59:50.700 --> 01:59:55.624
But the agents were able to use them in an unintended way to compromise this other company.

2046
01:59:55.904 --> 01:59:56.825
We can't trust the agents.

2047
01:59:57.746 --> 01:59:58.466
We can't trust the agents.

2048
01:59:58.486 --> 02:00:00.148
We can trust them to be clever.

2049
02:00:00.728 --> 02:00:01.989
Yeah, to be very, very clever.

2050
02:00:03.771 --> 02:00:04.832
What are your closing remarks?

2051
02:00:05.312 --> 02:00:08.154
You know, to the people that are listening right now, we've talked about lots of things.

2052
02:00:08.275 --> 02:00:09.576
Where is the right place to close?

2053
02:00:10.677 --> 02:00:12.278
What is your conclusive statement?

2054
02:00:13.818 --> 02:00:15.819
I just got married in July.

2055
02:00:16.319 --> 02:00:16.720
Congrats.

2056
02:00:17.440 --> 02:00:18.861
I'm the luckiest man in the world.

2057
02:00:20.101 --> 02:00:23.603
I have a mix of dread and excitement about the future.

2058
02:00:23.623 --> 02:00:27.444
I, like, really want us to make it through.

2059
02:00:29.425 --> 02:00:35.108
And so I'm just working really hard to try to help us figure it out.

2060
02:00:36.229 --> 02:00:41.093
we can fight all day long about, you know, who should be first and how it should all work.

2061
02:00:41.794 --> 02:00:44.556
But at the end of the day, we are facing this common threat.

2062
02:00:44.636 --> 02:00:45.337
We really are.

2063
02:00:46.518 --> 02:00:46.698
And,

2064
02:00:48.241 --> 02:00:49.341
I want people's help with that.

2065
02:00:50.062 --> 02:00:51.142
I don't think it works.

2066
02:00:51.762 --> 02:00:59.685
If we all just sit around and we like are very, you know, around social media all the time and that's just all we're doing, like, okay, companies will make more and more powerful AIs.

2067
02:01:00.426 --> 02:01:01.426
They'll make more and more money.

2068
02:01:02.426 --> 02:01:06.628
And eventually they build super intelligence and we lose, whether it's the U.S. or China.

2069
02:01:08.509 --> 02:01:09.469
We don't have to do that.

2070
02:01:11.713 --> 02:01:15.521
And I think people often feel like it's too big.

2071
02:01:16.422 --> 02:01:17.404
It's like too large.

2072
02:01:17.484 --> 02:01:22.193
It's like these giant, you know, multi-billion dollar corporations, this geopolitics.

2073
02:01:22.394 --> 02:01:22.975
We feel small.

2074
02:01:22.995 --> 02:01:23.837
We feel disempowered.

2075
02:01:25.750 --> 02:01:30.032
And I actually think that this is an area where people can do a lot.

2076
02:01:30.872 --> 02:01:34.534
Like, I actually think that people can help quite a bit.

2077
02:01:35.274 --> 02:01:38.616
And the reason I know this is because I've been going and talking to members of Congress.

2078
02:01:39.096 --> 02:01:40.117
I've talked with Bernie Sanders.

2079
02:01:40.537 --> 02:01:43.178
I've talked with, like, a bunch of senators on both the left and the right.

2080
02:01:43.918 --> 02:01:46.980
And they are starting to realize that this is very different.

2081
02:01:47.180 --> 02:01:51.422
And this is something that's happening that could really threaten our safety.

2082
02:01:52.102 --> 02:01:55.024
The closing question left from the last guest kind of links to this, so I'll ask it now.

2083
02:01:55.324 --> 02:01:55.504
Yes.

2084
02:01:55.524 --> 02:01:59.487
What is a simple thing the audience could do to create a better future?

2085
02:01:59.927 --> 02:02:04.050
So one of the things that works if enough people do it is calling your representative.

2086
02:02:04.990 --> 02:02:10.133
So some of my friends made a site, callcongress.ai, that walks you through exactly how to do it.

2087
02:02:10.894 --> 02:02:14.136
I think sometimes it seems like a little cheesy or a little bit like, that doesn't really work, right?

2088
02:02:14.376 --> 02:02:15.557
I'm like, no, it actually does work.

2089
02:02:15.917 --> 02:02:16.838
I have talked to these people.

2090
02:02:17.178 --> 02:02:20.700
And if their constituents come to them and say they're very worried about this, like

2091
02:02:20.840 --> 02:02:21.681
They have to get re-elected.

2092
02:02:22.001 --> 02:02:24.424
And they're also starting to get concerned themselves.

2093
02:02:24.824 --> 02:02:31.050
If they see a signal from their constituents that this is a very important issue to them, I think Congress can act.

2094
02:02:31.230 --> 02:02:36.556
I actually think that's also much of the solution here.

2095
02:02:37.617 --> 02:02:37.937
Power.

2096
02:02:39.136 --> 02:02:41.139
is driving motivations in one direction at the moment.

2097
02:02:41.239 --> 02:02:45.124
But staying in power from a political standpoint is also a pretty powerful incentive.

2098
02:02:45.524 --> 02:02:52.153
And as we think about 2028, the election cycle, I think AI is going to be one of the most important subjects on the ballot.

2099
02:02:52.733 --> 02:02:53.574
And the electorate,

2100
02:02:54.607 --> 02:02:56.508
really are aligned in what they want to hear.

2101
02:02:57.008 --> 02:02:57.968
They want their jobs preserved.

2102
02:02:57.988 --> 02:02:58.848
They want safety.

2103
02:02:59.488 --> 02:03:00.889
They want a future for their children.

2104
02:03:02.149 --> 02:03:05.750
So Trump, for example, I know he can't be reelected legally.

2105
02:03:06.510 --> 02:03:10.391
If he could get a third term, I think he would have to change his position to get elected in 2028.

2106
02:03:10.451 --> 02:03:10.831
Yeah.

2107
02:03:11.092 --> 02:03:12.732
Incentives aren't just a thing that happen out there.

2108
02:03:12.752 --> 02:03:13.972
Like, we are part of the incentives.

2109
02:03:14.092 --> 02:03:14.232
Yeah.

2110
02:03:14.272 --> 02:03:15.273
We provide the incentives.

2111
02:03:15.553 --> 02:03:15.993
Yeah, for now.

2112
02:03:16.253 --> 02:03:16.953
Yeah, for now.

2113
02:03:17.833 --> 02:03:18.414
Jeffrey, thank you.

2114
02:03:18.854 --> 02:03:19.254
Yeah, thank you.

2115
02:03:19.274 --> 02:03:19.754
Thank you so much.
