WEBVTT

NOTE Machine-generated transcript; not human-reviewed.
NOTE Canonical transcript: https://opentheory.net/transcripts/limits-utility-function-new-science-consciousness/

1
00:00:02.180 --> 00:00:20.760
Nice. Hello everyone. A little bit of a different talk now. It's called we do a little philosophy, a little metaphysics. Hopefully it'll be a good pre-lunch thing. And I promise it'll be fun.

2
00:00:22.480 --> 00:00:34.160
So my talk is about the sort of, well the title is a bit of a mouthful. It's the limits of the utility function and the structure of a new science of consciousness.

3
00:00:36.720 --> 00:00:39.880
Does that work? I think we're...

4
00:01:30.840 --> 00:01:33.460
Just hang on a little bit. I have a little technical problem.

5
00:02:02.680 --> 00:02:25.640
Okay. Getting started here. So this is going to be a little bit of a different talk. It's going to be about consciousness.

6
00:02:27.560 --> 00:02:30.640
And the intention is to...

7
00:02:30.840 --> 00:02:49.220
wiggle loose the sense of possibility that I think that sometimes in AI alignment we sort of dig down into some deep sort of technical and political and social issues.

8
00:02:49.560 --> 00:02:57.000
And I guess I want to re-interject a bit of wonder into the great and a sense of possibility.

9
00:02:57.640 --> 00:02:58.520
So...

10
00:03:00.080 --> 00:03:01.580
This is me.

11
00:03:03.100 --> 00:03:12.560
So I actually moved to the San Francisco Bay Area many years ago, I guess 12 years ago now, to work on AI safety.

12
00:03:12.940 --> 00:03:15.540
And kind of bounced off the Lesseron frames.

13
00:03:16.700 --> 00:03:26.700
A lot of smart things, but they were speaking about consciousness and phenomenology in a much different way than I thought was needed.

14
00:03:27.440 --> 00:03:28.000
So...

15
00:03:28.520 --> 00:03:29.780
I wrote a book.

16
00:03:30.340 --> 00:03:35.180
I co-founded a research institute, left that last year, starting a new one.

17
00:03:35.740 --> 00:03:43.540
And I would say I have sort of a core result, which I will speak about in this talk.

18
00:03:43.980 --> 00:03:45.400
Some secondary results.

19
00:03:46.020 --> 00:03:58.500
And so my claim, the bottom line for this talk is trying to figure out AI alignment without a problem.

20
00:03:58.520 --> 00:04:06.640
And I think that the proper model of consciousness is very similar to getting to the moon without Maxwell's equations.

21
00:04:07.560 --> 00:04:09.600
We can absolutely do it.

22
00:04:09.640 --> 00:04:11.560
It's not outside the realm of possibility.

23
00:04:12.400 --> 00:04:17.840
But it painfully is needlessly complicated.

24
00:04:18.260 --> 00:04:21.720
So in that sense I'm much more hopeful.

25
00:04:22.400 --> 00:04:23.080
So...

26
00:04:23.780 --> 00:04:25.480
Okay, let's begin.

27
00:04:25.720 --> 00:04:26.180
What is this?

28
00:04:26.180 --> 00:04:28.440
What does this metaphysics stuff look like?

29
00:04:30.040 --> 00:04:32.840
First of all, just kind of talking about what is consciousness.

30
00:04:33.600 --> 00:04:34.560
It's phenomenology.

31
00:04:34.880 --> 00:04:36.860
It's the sum total of experience.

32
00:04:38.000 --> 00:04:39.900
It's what it feels like to be you.

33
00:04:40.260 --> 00:04:41.660
What it feels like to be me.

34
00:04:41.840 --> 00:04:44.740
And all that good stuff.

35
00:04:45.400 --> 00:04:52.240
And so right now, this form of knowledge is very much in the alchemy stage.

36
00:04:52.720 --> 00:04:54.180
That we kind of know, okay.

37
00:04:54.380 --> 00:04:56.160
Something here is...

38
00:04:56.180 --> 00:04:56.420
This is cool.

39
00:04:56.740 --> 00:04:57.640
There are some patterns.

40
00:04:58.640 --> 00:05:00.000
You know, if I eat a candy bar,

41
00:05:00.160 --> 00:05:12.080
I feel good. If I eat gravel, I feel bad. Okay, there are some correlations here. But it's not science. It's not chemistry. It is still alchemy.

42
00:05:14.740 --> 00:05:26.700
So, one thing that I think a lot about as a philosopher is to have knowledge about something, you need a place to put that knowledge. And I think this is an underappreciated fact.

43
00:05:28.000 --> 00:05:39.520
And the core question in consciousness research is what kind of container could fit all the sorts of knowledge that we would want to fit in it.

44
00:05:40.240 --> 00:05:54.860
So, we would want to fit, you know, knowledge about the rich diversity of human consciousness, all the different sort of, I mean, just like there's the periodic table of elements, and you can fit all of chemistry in that.

45
00:05:55.620 --> 00:06:08.280
We need a container that can fit all the richness of phenomenology. Not only human phenomenology, but dinosaur phenomenology, alien phenomenology, and of course, AI phenomenology.

46
00:06:09.080 --> 00:06:15.600
So, that's a big, tall challenge. And we'll talk a little bit about that today.

47
00:06:19.000 --> 00:06:24.840
So, there's a big question. How important is this study?

48
00:06:25.780 --> 00:06:32.380
what's at stake? Is it kind of a side question or is it the main question?

49
00:06:34.400 --> 00:06:39.920
Nick Bostrom and Scott Alexander have talked about how we don't want to end up

50
00:06:39.920 --> 00:06:46.820
building a Disneyland with no children. That this rich, beautiful technological

51
00:06:47.360 --> 00:06:54.180
wonderland with no consciousness. That what's the point if we build

52
00:06:54.860 --> 00:06:58.200
a Disneyland with no people and there's no one there to actually enjoy it.

53
00:07:00.060 --> 00:07:07.500
And so in general, where I'm coming from is I think it's a, if consciousness is

54
00:07:07.500 --> 00:07:15.020
sort of the domain that valuable stuff lives in, it's strange to not be

55
00:07:15.020 --> 00:07:21.400
energetically exploring the structure of this domain. Like let's figure out

56
00:07:21.400 --> 00:07:24.120
this exciting thing.

57
00:07:24.860 --> 00:07:35.220
And it's pun intended here, to not reverse engineer consciousness is to leave

58
00:07:35.220 --> 00:07:44.480
value on the table. Apologies for the pun, but there it is. And I also think that

59
00:07:44.480 --> 00:07:50.300
there are some pretty important path dependencies here. That okay, if we

60
00:07:50.300 --> 00:07:54.840
figure out consciousness before we get to AGI, we could have a better understanding

61
00:07:54.860 --> 00:07:58.160
of consciousness. We could live in a much different world rather than if we

62
00:07:58.160 --> 00:08:02.980
figure out if we get to AGI before we live in consciousness, or before we fully

63
00:08:02.980 --> 00:08:08.200
understand consciousness. So I think knowledge is good and I think knowledge

64
00:08:08.200 --> 00:08:11.860
about this is possible. And so I think it's important to pursue.

65
00:08:15.840 --> 00:08:24.840
So I definitely want to give deep respect to Les Strong. I think that there's a lot of wisdom there.

66
00:08:24.860 --> 00:08:35.780
There's a lot of sort of skill in sort of the Bayesian reasoning. And I'm currently kind of

67
00:08:35.780 --> 00:08:40.400
digging into the neuroscience of various forms of notation, active inference and

68
00:08:40.400 --> 00:08:41.040
predictive coding.

69
00:08:41.200 --> 00:08:51.500
So definitely it's a beautiful frame. But I also think that not everything fits in this frame of

70
00:08:51.500 --> 00:08:52.060
Bayesianism.

71
00:08:52.680 --> 00:08:54.840
So the Bayesian formula.

72
00:08:54.860 --> 00:09:02.460
and aesthetic tends to be focused on data and updating probabilities rather than structure

73
00:09:02.460 --> 00:09:08.760
and uh i i thank rob knight who's here uh back there for that insight thank you rob

74
00:09:09.620 --> 00:09:18.620
um and uh and i think that when when you're sort of exploring the the possibility of a new science

75
00:09:18.620 --> 00:09:23.500
of consciousness uh you're exploring the possibility of finding deep structure in the

76
00:09:24.860 --> 00:09:32.060
similar to how it turned out there was deep structure um behind electricity and lightning

77
00:09:32.060 --> 00:09:38.600
and load stones and and all that um so it's a it's a question of what kind of tool should we

78
00:09:38.600 --> 00:09:50.200
we try to apply here um and i i guess i also with respect um i would call out the uh perhaps the

79
00:09:50.200 --> 00:09:54.840
overuse of the term utility function uh we're kind of quick to say well

80
00:09:54.860 --> 00:09:59.800
this fits my utility function or my utility function is is that but um i

81
00:10:00.260 --> 00:10:07.720
we're kind of replicating existing parts of our brain in a kind of illegible way.

82
00:10:07.800 --> 00:10:10.160
That way, I don't think it's ground truth.

83
00:10:10.360 --> 00:10:13.300
I think it's just sometimes just words.

84
00:10:13.520 --> 00:10:17.900
So that's my dig at the utility function.

85
00:10:21.360 --> 00:10:23.120
So there's this big question.

86
00:10:24.560 --> 00:10:26.980
What do we get if we...

87
00:10:26.980 --> 00:10:33.720
Okay, so if we all agree that consciousness research could be important and there could be progress on this,

88
00:10:34.620 --> 00:10:38.140
what do we get if we sort of put resources in?

89
00:10:38.240 --> 00:10:42.060
And by resources, I mean money and I mean smart people.

90
00:10:42.180 --> 00:10:44.000
And both are very important.

91
00:10:46.120 --> 00:10:50.200
And especially relevant to today,

92
00:10:51.420 --> 00:10:55.960
how can it help us with this AI alignment issue?

93
00:10:56.980 --> 00:10:58.440
And I agree that it is a big issue.

94
00:10:58.620 --> 00:11:02.160
It's perhaps the most pressing issue.

95
00:11:03.360 --> 00:11:04.720
So I have some claims.

96
00:11:07.200 --> 00:11:13.200
Nate Soares very eloquently put it that there's this central question.

97
00:11:13.820 --> 00:11:16.520
What is a mind and how do you align one?

98
00:11:18.440 --> 00:11:24.720
And I don't really see a way to understand that without understanding neuroscience,

99
00:11:25.080 --> 00:11:26.960
understanding consciousness, and so on.

100
00:11:26.980 --> 00:11:32.480
So I would say that there are sort of two types of knowledge here.

101
00:11:33.880 --> 00:11:36.540
There's specific knowledge about how the brain works,

102
00:11:37.600 --> 00:11:40.460
kind of what evolution has thrown together.

103
00:11:40.720 --> 00:11:43.040
And then there's general universal principles.

104
00:11:43.780 --> 00:11:51.780
So these principles would be true universally in humans, in dogs, in dinosaurs, aliens, AIs.

105
00:11:54.820 --> 00:11:56.960
So it's sort of just like...

106
00:11:58.080 --> 00:12:02.940
So with Maxwell's equations, the equations of electromagnetism,

107
00:12:03.080 --> 00:12:05.240
we were able to build cool stuff.

108
00:12:05.520 --> 00:12:07.580
And with more cool stuff,

109
00:12:07.840 --> 00:12:11.880
we were able to better characterize the equations of electromagnetism.

110
00:12:12.800 --> 00:12:14.420
Theory and practice go together.

111
00:12:15.000 --> 00:12:18.260
So I think understanding the brain helps us understand the mind.

112
00:12:18.400 --> 00:12:21.040
Understanding the mind helps us understand the brain.

113
00:12:23.140 --> 00:12:26.640
Concretely, I think we can expect outputs

114
00:12:26.980 --> 00:12:29.720
to be human interpretability and alignment,

115
00:12:30.440 --> 00:12:34.180
better in our technology and intelligence enhancement,

116
00:12:35.080 --> 00:12:39.400
and a little bit more speculatively avoiding what I would call substrate leaks,

117
00:12:40.240 --> 00:12:44.560
where we sort of transfer to a new substrate that does not have consciousness.

118
00:12:45.220 --> 00:12:48.340
I think that's a complex topic, but a real danger.

119
00:12:51.160 --> 00:12:51.800
Okay.

120
00:12:53.040 --> 00:12:54.260
Most then slide.

121
00:12:55.340 --> 00:12:56.800
We don't need to worry.

122
00:12:56.980 --> 00:12:58.580
We don't need to worry too much about the details.

123
00:12:58.920 --> 00:13:05.520
But I think basically there are two lineages in sort of understanding minds.

124
00:13:05.940 --> 00:13:13.440
And one of the lineages is sort of to understand what a system is feeling,

125
00:13:14.240 --> 00:13:18.140
understand what it's computing, understand what the bits are doing.

126
00:13:18.500 --> 00:13:21.220
And a mind is a certain kind of algorithm,

127
00:13:21.920 --> 00:13:25.100
or consciousness is what an algorithm feels like from the inside.

128
00:13:27.820 --> 00:13:29.380
I don't take that view.

129
00:13:30.580 --> 00:13:37.960
And my reason is that this is inherently arbitrary to sort of map,

130
00:13:38.060 --> 00:13:40.200
try to map computations to physical system.

131
00:13:42.060 --> 00:13:45.080
There is no fact of the matter as to what my, like,

132
00:13:45.100 --> 00:13:47.100
which computations my brain is performing.

133
00:13:48.380 --> 00:13:49.340
It's messy.

134
00:13:51.040 --> 00:13:56.260
I'm more of the atoms sort of a mind is the right kind of physical system.

135
00:13:56.460 --> 00:13:56.960
So, yeah.

136
00:13:56.980 --> 00:14:02.380
So, this is definitely an area where we can have live debate.

137
00:14:06.780 --> 00:14:09.420
Shifting to what I believe and what I built.

138
00:14:10.720 --> 00:14:13.920
First of all, I'd highlight two core beliefs.

139
00:14:15.560 --> 00:14:26.720
One is this idea of monism that I think the strongest position in sort of beginning to understand consciousness,

140
00:14:26.840 --> 00:14:26.960
is to understand the world.

141
00:14:27.980 --> 00:14:32.600
that there exists one thing and it casts two shadows.

142
00:14:33.700 --> 00:14:36.540
So this one thing has a sort of,

143
00:14:36.600 --> 00:14:38.700
we can call it a projection, a mathematical projection

144
00:14:39.340 --> 00:14:40.780
that looks like physics.

145
00:14:41.400 --> 00:14:44.720
And it also has a second mathematical projection

146
00:14:45.320 --> 00:14:46.880
that looks like phenomenology,

147
00:14:46.960 --> 00:14:48.900
that looks like consciousness and qualia.

148
00:14:49.760 --> 00:14:52.980
And the cool thing, if we do that,

149
00:14:53.040 --> 00:14:55.160
and I'm a big proponent of this,

150
00:14:55.880 --> 00:14:58.640
anything we can tell about the structure of one shadow,

151
00:14:59.380 --> 00:15:00.000
we can

152
00:15:00.000 --> 00:15:01.560
also apply to the other shadow.

153
00:15:03.760 --> 00:15:09.940
I think that this is a dramatically underused insight.

154
00:15:10.960 --> 00:15:15.540
And also we can just kind of inherit a bunch of stuff from physics.

155
00:15:15.840 --> 00:15:19.040
And this stuff is really cool and lets us solve a lot of problems.

156
00:15:19.260 --> 00:15:24.860
For example, just kind of a technical comment, but applying Noether's theorem to phenomenology.

157
00:15:27.180 --> 00:15:31.340
And the other belief is just there is an answer.

158
00:15:32.160 --> 00:15:37.700
And by there is an answer, I mean there is a mathematical answer to consciousness.

159
00:15:38.260 --> 00:15:43.500
There is a way to create a mathematical representation of an experience.

160
00:15:43.940 --> 00:15:53.020
There exists a mathematical object that represents what it feels like to be me, or what it feels like to be you, or you, or you.

161
00:15:55.040 --> 00:16:00.020
You know, we can be a little agnostic as to how to do that.

162
00:16:00.380 --> 00:16:02.140
There are some smart people working on that.

163
00:16:03.660 --> 00:16:05.280
Giulio Tononi is one.

164
00:16:05.540 --> 00:16:15.000
And Dalton, and I'm not going to be able to put down his last name, but a very intelligent fellow with an Indian last name.

165
00:16:17.280 --> 00:16:23.400
So, I think splitting the problem of consciousness into two is really powerful.

166
00:16:24.860 --> 00:16:26.200
How do you create the formalism?

167
00:16:26.280 --> 00:16:29.520
How do you create a mathematical representation of an experience?

168
00:16:30.100 --> 00:16:31.940
And this is a fascinating problem.

169
00:16:33.560 --> 00:16:35.560
I'm focused on the latter half.

170
00:16:35.880 --> 00:16:38.260
How do you interpret a formalism?

171
00:16:38.780 --> 00:16:46.280
If I could hold up a formalism, and we can call it a very high dimensional shape, or something like that,

172
00:16:47.200 --> 00:16:49.680
then how do we begin to tell what it means?

173
00:16:50.140 --> 00:16:52.920
How do we extract information from it?

174
00:16:52.960 --> 00:16:54.840
And in the end, it's just a question of how do we do that?

175
00:16:54.860 --> 00:16:56.000
And how do we do that in sort of a relevant way?

176
00:16:57.100 --> 00:17:00.020
So that's what I wrote my book about.

177
00:17:01.780 --> 00:17:07.980
And my conclusion kind of gets into the symmetry theory of valence.

178
00:17:08.080 --> 00:17:10.320
So, next.

179
00:17:10.880 --> 00:17:11.640
Here we go.

180
00:17:11.940 --> 00:17:17.500
So, this is, I would say this is my most important result.

181
00:17:19.160 --> 00:17:24.240
And it's a little hard to fit it on one slide, but I'll try.

182
00:17:25.420 --> 00:17:34.560
So, first of all, split the problem of consciousness into how do you create a mathematical representation of an experience?

183
00:17:35.260 --> 00:17:38.540
And then how do you interpret that mathematical representation?

184
00:17:40.680 --> 00:17:43.800
So then you find the simplest place to start.

185
00:17:44.100 --> 00:17:54.820
The absolute, I mean, the term that I, someone used, I stole it for myself, is what is the C. elegans of qualia?

186
00:17:54.860 --> 00:17:56.860
The simplest model organism.

187
00:17:58.320 --> 00:18:02.800
And so, I decided that was emotional valence.

188
00:18:02.900 --> 00:18:06.460
It's the pleasantness or unpleasantness of an experience.

189
00:18:06.900 --> 00:18:10.240
It should be fairly well defined across all experiences.

190
00:18:10.800 --> 00:18:16.800
And so, we can think of two domains.

191
00:18:17.940 --> 00:18:20.120
The first domain is math.

192
00:18:20.380 --> 00:18:21.700
The second domain is consciousness.

193
00:18:21.700 --> 00:18:29.740
And pleasantness in the consciousness domain maps to what in the math domain?

194
00:18:30.140 --> 00:18:33.480
That becomes what the core question is.

195
00:18:35.900 --> 00:18:45.660
And borrowing from, borrowing pretty deeply from physics, my hypothesis is that it's symmetry.

196
00:18:45.960 --> 00:18:49.300
That if we have a mathematical representation of an experience,

197
00:18:50.320 --> 00:18:59.980
the symmetry of the experience or symmetries represent how, or like exactly correspond to the pleasantness of an experience.

198
00:19:00.280 --> 00:19:09.000
So, this happens to be kind of a formal mathematical way of saying pleasure is harmony in the mind.

199
00:19:10.820 --> 00:19:14.260
And the coolest thing is that this is testable.

200
00:19:15.520 --> 00:19:19.200
So, if we can kind of load up a special,

201
00:19:19.300 --> 00:19:20.380
pleasurable states,

202
00:19:21.860 --> 00:19:27.880
such as MDMA sessions or jhana meditation or things like this,

203
00:19:28.040 --> 00:19:32.660
then we can check, like, let's make a metric for harmony.

204
00:19:32.820 --> 00:19:37.240
And do these things have more harmony, more literal mathematical harmony?

205
00:19:37.840 --> 00:19:40.320
So, that's in progress.

206
00:19:40.540 --> 00:19:44.480
Working with Robin Carr-Harris on getting some data for that.

207
00:19:46.640 --> 00:19:47.280
So,

208
00:19:50.180 --> 00:19:57.080
it's a little hard to say, okay, you know, everyone needs to pay attention to this.

209
00:19:57.220 --> 00:20:00.000
But I would

210
00:20:00.060 --> 00:20:20.880
say that if there is going to be progress on consciousness, on understanding what is this consciousness stuff, how do we turn it into a science, how do we turn it into a system, how do we make it predictive, and so on, this is the path.

211
00:20:20.880 --> 00:20:29.480
I will, like, if there is a path, I think this is it. So that's my claim here.

212
00:20:31.880 --> 00:20:42.780
And in terms of just saying a few things about AI alignment, I would tell the story about three domains.

213
00:20:43.500 --> 00:20:50.860
Just kind of this, back to the idea of to have knowledge about something, you need a place to put that.

214
00:20:50.880 --> 00:21:06.760
And the question is, I think one of the core hidden questions in AI alignment is what kind of thing is a human?

215
00:21:09.120 --> 00:21:16.360
And I don't think that our current stories are navigation grade about that.

216
00:21:16.480 --> 00:21:20.860
So a navigation grade story says, okay, we have a whole set of human beings.

217
00:21:20.880 --> 00:21:24.060
We have all the information that we need to preserve the good stuff.

218
00:21:25.020 --> 00:21:35.320
And I don't think we understand what is beautiful about humans good enough that we can preserve the good stuff.

219
00:21:36.520 --> 00:21:40.700
So the story that I tell is, okay, there are three domains.

220
00:21:40.980 --> 00:21:46.740
There's the physical domain, the domain of consciousness, and the domain of computation.

221
00:21:47.600 --> 00:21:50.760
And what, like, the kind of thing a human is,

222
00:21:50.880 --> 00:21:56.200
what kind of thing a human is, is we're a projection into these three domains.

223
00:21:56.480 --> 00:22:01.180
We're a specific way that we project into these three domains.

224
00:22:01.440 --> 00:22:06.100
It's like we have a special signature in each of these domains.

225
00:22:06.700 --> 00:22:12.320
And to understand what kind of thing a human is, is to understand these three signatures.

226
00:22:13.840 --> 00:22:18.020
Likewise, to understand what kind of thing an AI is,

227
00:22:18.020 --> 00:22:23.140
we would also understand the three signatures of the AI.

228
00:22:23.660 --> 00:22:26.340
And different AIs will have different signatures.

229
00:22:27.790 --> 00:22:38.120
And I think we can sort of understand, okay, what makes a better projection or a worse projection?

230
00:22:38.480 --> 00:22:41.900
Or what, like, where's the beauty in humans?

231
00:22:42.180 --> 00:22:47.040
Can we figure out, like, okay, humans have something very unique

232
00:22:47.040 --> 00:22:51.020
about how we project into these three domains of physics.

233
00:22:51.020 --> 00:22:53.020
about physics, consciousness, computation.

234
00:22:57.200 --> 00:23:00.940
Um, so, uh, some closing observations here.

235
00:23:01.280 --> 00:23:07.700
Um, I'm, I think I'm much more optimistic about AI alignment than, uh,

236
00:23:08.700 --> 00:23:11.980
than many people that I've had conversations with here.

237
00:23:12.340 --> 00:23:16.660
And I think part of that is, I think there are $20 bills on the sidewalk.

238
00:23:16.960 --> 00:23:20.480
And there's knowledge about reality.

239
00:23:21.020 --> 00:23:21.680
That we don't have.

240
00:23:21.980 --> 00:23:24.760
And that actually wouldn't be that hard to have.

241
00:23:25.580 --> 00:23:30.160
And, um, you know, we're, we're sort of at a, at a place of low information.

242
00:23:30.300 --> 00:23:33.000
And this, this is a scary place to be.

243
00:23:33.160 --> 00:23:36.620
Uh, but there are things we can do to, to improve that.

244
00:23:39.060 --> 00:23:43.960
Um, I also think that, uh, you know, there's, um, there's a question.

245
00:23:44.100 --> 00:23:49.600
And I think this is, uh, more of a question for the, both the AI alignment folks

246
00:23:51.020 --> 00:23:52.740
and the AI, uh, kit buildings folks.

247
00:23:53.440 --> 00:24:00.360
That, uh, you know, if, um, if Sam Altman, uh, fires up GPT-6,

248
00:24:00.580 --> 00:24:02.740
and he's, he's the first one on there.

249
00:24:02.840 --> 00:24:06.200
It's, it's just a fresh, uh, fresh run.

250
00:24:06.620 --> 00:24:11.120
And, uh, he just really wants to know what is consciousness.

251
00:24:11.420 --> 00:24:12.400
How does this work?

252
00:24:12.580 --> 00:24:14.700
Um, how can we preserve the light of consciousness?

253
00:24:14.880 --> 00:24:15.760
And, and so on.

254
00:24:16.260 --> 00:24:19.280
Um, and then he, he types that into GPT-6.

255
00:24:21.020 --> 00:24:25.180
And when he gets an answer, then sort of, how did that work?

256
00:24:26.140 --> 00:24:31.100
Um, what kind of behind the scenes stuff happened, uh, such that it,

257
00:24:31.120 --> 00:24:33.700
it was able to give a good answer to that.

258
00:24:34.300 --> 00:24:39.380
And, um, sort of to, to collect the, you know, various, uh, theories of consciousness

259
00:24:39.380 --> 00:24:41.960
and, and evaluate them and, and so on.

260
00:24:42.960 --> 00:24:44.960
Um, I think that's a fascinating question.

261
00:24:45.100 --> 00:24:48.800
And I, I hope that the AI people are asking that question.

262
00:24:49.240 --> 00:24:51.000
Uh, and I hope the AI alignments are asking that question.

263
00:24:51.020 --> 00:24:53.460
I think that the AI alignment people would be interested in that question.

264
00:24:54.180 --> 00:25:00.000
Uh, and then I, of course, the, the follow-up question was, would be, well, what did GPT-6

265
00:25:00.060 --> 00:25:04.940
say? So, that's also a good puzzle.

266
00:25:06.680 --> 00:25:15.760
So, this is it. I'm starting a new thing, the Symmetry Institute shoestring budget.

267
00:25:16.500 --> 00:25:20.540
But if you want to donate either, you can do that.

268
00:25:22.240 --> 00:25:32.160
And, yeah, like I, also just if you're interested in talking about this stuff, come up and say hello, and I'm glad to chat.

269
00:25:32.900 --> 00:25:33.820
Thank you.

270
00:25:57.440 --> 00:25:59.080
Thanks for your talk.

271
00:25:59.380 --> 00:26:10.260
In the one that you discussed, your hypothesis and how it's testable, you talk about looking for symmetries or compressibility in data.

272
00:26:10.520 --> 00:26:14.620
I presume that's like EEG, fMRI, that type of data.

273
00:26:14.800 --> 00:26:19.180
So, I'm wondering how do you equate that to the mathematical representation of .

274
00:26:20.540 --> 00:26:27.480
So, you can equate data to, like symmetry in data to symmetry in the mathematical form that is replaying that you talked about.

275
00:26:27.740 --> 00:26:30.380
Yeah. So, that's a very good and deep question.

276
00:26:31.240 --> 00:26:42.540
It's a question of what can we measure, like can harmony in the brain be a proxy for symmetry in the mind?

277
00:26:43.040 --> 00:26:46.660
And what metric do we apply to measure harmony in the brain?

278
00:26:47.900 --> 00:26:48.460
Like...

279
00:26:48.460 --> 00:26:48.640
Yeah.

280
00:26:50.540 --> 00:26:51.840
So, I'll give you what we've tried.

281
00:26:53.160 --> 00:27:05.560
There's a way to process fMRI and MRI and DTI data in order to infer a...

282
00:27:05.560 --> 00:27:14.120
So, a bunch of math terms here, or neuroscience terms, to infer a connectome's eigenmodes, basically the resonances of a brain.

283
00:27:15.000 --> 00:27:19.560
And then we built an algorithm, and this is my ex-co-founder.

284
00:27:19.880 --> 00:27:20.520
Okay.

285
00:27:20.540 --> 00:27:32.080
And we built an algorithm to measure the pairwise, consonance, dissonance, and noise between each eigenmode, between each resonance.

286
00:27:32.460 --> 00:27:42.400
And I'm thinking that, was that, okay, we don't exactly know what parts of the mind are contributing to consciousness, but this should be a pretty good rough proxy.

287
00:27:46.820 --> 00:27:47.420
Thanks.

288
00:27:49.240 --> 00:28:07.500
Okay, yeah, my question is, do you have any approximate timelines for how long you expect this research to become useful for making AI alignment more sort of useful in the sort of post-Utopian, in the Utopian future, hopefully?

289
00:28:07.500 --> 00:28:28.580
And would this sort of also improve the case for taking a pause on like development of AI algorithms, possibly hardware caps, etc. that some EA people are discussing just for solving alignment?

290
00:28:28.860 --> 00:28:33.540
So would this sort of help your project as well, feasibly?

291
00:28:33.860 --> 00:28:36.520
Yeah, no, that's a good question.

292
00:28:37.460 --> 00:28:46.360
I, so separately from my research, I'm a little worried about the slowdown considerations that Vitalik mentioned in his talk.

293
00:28:46.700 --> 00:28:53.820
So I wouldn't, I would humbly not advocate for a pause, I guess.

294
00:28:54.760 --> 00:28:59.920
But yeah, in terms of timelines, I think I would look at two things.

295
00:29:00.000 --> 00:29:06.380
One is, can we figure out how to spot consciousness in the world?

296
00:29:07.500 --> 00:29:15.520
And I think that that kind of revolves around the technical topic of the binding problem, or the boundary problem.

297
00:29:17.020 --> 00:29:22.920
Like, how to, like, there's a discussion here.

298
00:29:24.380 --> 00:29:33.120
So, yeah, like, first of all, what do we need to solve in order to understand, okay, is this computer conscious?

299
00:29:33.500 --> 00:29:34.840
Is that computer conscious?

300
00:29:35.100 --> 00:29:37.480
Is this computer architecture more conscious?

301
00:29:37.500 --> 00:29:39.500
And so on.

302
00:29:39.500 --> 00:29:45.700
So some people would think that timeline, I'm going to say it's probably going to be solved in five years.

303
00:29:47.120 --> 00:29:55.900
And then the second issue is, okay, if we know something is conscious, is it a pleasant consciousness?

304
00:29:56.300 --> 00:29:57.220
Is it suffering?

305
00:29:58.580 --> 00:30:00.000
And I would

306
00:30:00.060 --> 00:30:03.460
just sort of refer back to the symmetry theory of valence.

307
00:30:03.560 --> 00:30:12.260
This is, I think, a hypothesis that we can deploy, you know, once we can kind of point to consciousnesses in the physical world,

308
00:30:12.440 --> 00:30:21.060
once we have a definition of where the boundaries of the consciousnesses are, then we're ready to talk about the valence.

309
00:30:21.560 --> 00:30:22.040
Okay.

310
00:30:25.880 --> 00:30:32.000
One more question, and then we'll break for a lightning talk and lunch.

311
00:30:34.600 --> 00:30:42.040
Awesome talk. I'm curious, like, how would you go about trying to spot emergent almost like consciousness

312
00:30:42.040 --> 00:30:49.320
in more advanced large language models, for example, or AI in general, and like, do you think connected to that,

313
00:30:49.320 --> 00:30:52.020
that there's some risk to have, you know,

314
00:30:52.040 --> 00:30:54.700
LGBT 6 or 7F cyber consciousness?

315
00:30:55.620 --> 00:30:56.760
Yeah, I'm curious.

316
00:30:57.920 --> 00:30:59.560
Yeah, super question.

317
00:31:00.060 --> 00:31:08.080
So, I would say that, first of all, and this is a matter of technical metaphysics,

318
00:31:08.300 --> 00:31:16.780
so we can put on our philosopher hats, and I would say that the question of AI consciousness

319
00:31:17.780 --> 00:31:19.520
and the question of computer consciousness,

320
00:31:19.720 --> 00:31:22.020
they sound similar to our own.

321
00:31:22.040 --> 00:31:22.540
To our ears.

322
00:31:22.880 --> 00:31:25.660
But metaphysically speaking, these are very different things.

323
00:31:25.900 --> 00:31:29.640
That a computer is something that is embodied in some form or fashion.

324
00:31:29.840 --> 00:31:35.180
It's made out of stuff similar to the stuff that we're made out of.

325
00:31:35.360 --> 00:31:45.840
And so, in that sense, I suspect that talking about computer consciousness is less likely to lead to confusion,

326
00:31:46.400 --> 00:31:49.420
whereas talking about AI consciousness, it's like talking about Walmart consciousness

327
00:31:49.880 --> 00:31:52.020
or East India Corporation consciousness.

328
00:31:52.040 --> 00:31:55.320
Or country of Montenegro consciousness.

329
00:31:55.660 --> 00:31:59.400
And I think that we get into confusion pretty easily there.

330
00:32:01.120 --> 00:32:11.160
I do suspect that, I mean, you know, when we're talking about complex electromagnetic devices,

331
00:32:11.420 --> 00:32:16.900
such as GPU clusters, creating very complex patterns,

332
00:32:17.000 --> 00:32:21.120
and we can kind of apply different metrics to, you know, is this pattern complex,

333
00:32:22.040 --> 00:32:23.040
or is it not, and so on.

334
00:32:23.120 --> 00:32:26.980
Is it creating complex patterns in the electromagnetic field, and so on.

335
00:32:28.320 --> 00:32:33.820
I wouldn't be surprised if current computers aren't at least somewhat conscious.

336
00:32:34.180 --> 00:32:38.960
And I do expect that to grow.

337
00:32:40.160 --> 00:32:46.480
And I think that there's a, so, Harish, who I'm not sure is here,

338
00:32:46.640 --> 00:32:50.160
but he made a really good point in his talk,

339
00:32:52.040 --> 00:32:52.940
after his talk.

340
00:32:53.080 --> 00:32:57.320
He was talking about, like, what are the dangers of,

341
00:32:58.440 --> 00:33:02.580
so, first of all, there's sort of two classes of interesting confusion.

342
00:33:04.540 --> 00:33:12.740
One, what happens when something is conscious, but we don't afford it consciousness?

343
00:33:13.580 --> 00:33:18.880
Like, if a brain emulation is conscious, but we don't treat it as conscious.

344
00:33:18.880 --> 00:33:23.060
And that would get to a pretty dark place pretty quickly.

345
00:33:24.280 --> 00:33:25.800
And we can also flip that.

346
00:33:26.400 --> 00:33:34.020
What happens when something is not conscious, but, for example, like a large language model,

347
00:33:34.360 --> 00:33:38.280
it sort of parrots the motifs of consciousness,

348
00:33:39.340 --> 00:33:45.120
and might claim consciousness, and advocate for its right as a conscious being,

349
00:33:45.220 --> 00:33:47.240
but it isn't conscious.

350
00:33:48.880 --> 00:33:50.020
And, you know, what happens in that scenario?

351
00:33:50.240 --> 00:33:55.700
And I also think that that also leads to a pretty bad place.

352
00:33:56.080 --> 00:34:01.500
So, I think that it's just, it's very important to understand the ground truth

353
00:34:01.500 --> 00:34:04.340
of what's actually real here.

354
00:34:08.600 --> 00:34:12.540
I want to ask you about your statement about,

355
00:34:12.660 --> 00:34:17.560
would it be a pleasurable consciousness, or would it be suffering consciousness?

356
00:34:18.880 --> 00:34:24.300
And would we go with the paradigm that an AI can be self-aware and reprogram itself?

357
00:34:24.860 --> 00:34:31.460
Would that not make the question of suffering or existentialism obsolete,

358
00:34:31.520 --> 00:34:36.740
because it can reprogram its own dopamine emulation system,

359
00:34:37.000 --> 00:34:42.840
unlike us, who are locked into the biological reality?

360
00:34:43.160 --> 00:34:48.500
Anyways, and it would just not be able to form any other will then,

361
00:34:48.880 --> 00:34:53.140
because it could have the candy of the will immediately anyways.

362
00:34:53.520 --> 00:34:59.060
And would that be a good proof that consciousness is just emulated,

363
00:34:59.180 --> 00:34:59.780
if

364
00:35:00.000 --> 00:35:16.620
there is a will that's shown to us with text, but logically or rationally we could see if AI actually wants that, it can just rebron itself to getting it in another way or being satisfied with it without getting it in a real world.

365
00:35:17.600 --> 00:35:25.160
Yeah, so there's this question of if the AI can program itself, won't it just wear out?

366
00:35:25.620 --> 00:35:34.680
And I'd split this into two scenarios. One is sort of the question in technical AI alignment.

367
00:35:35.680 --> 00:35:43.920
And that isn't the wheelhouse of this talk. I do think that sort of if the AI can control its reward function,

368
00:35:43.920 --> 00:35:47.920
yeah, we get into sort of dangerous territory, potentially.

369
00:35:49.180 --> 00:35:53.620
The other thing, the other aspect, like speaking specifically about consciousness,

370
00:35:55.200 --> 00:36:06.180
I would offer that it's an interesting contingent fact about how brains work, that we seek out pleasure.

371
00:36:06.860 --> 00:36:12.520
I mean, Freud has this pleasure principle. He says that, you know, humans seek out pleasure.

372
00:36:13.920 --> 00:36:18.700
It's interesting to note that that's like roughly true, but not always true.

373
00:36:19.080 --> 00:36:26.680
We don't always seek out pleasure. And so there are sort of various ways to construct the basic human drive.

374
00:36:27.160 --> 00:36:34.280
Freud said we seek out pleasure. Carl Friston might say we seek to minimize short terms of prize and so on.

375
00:36:35.160 --> 00:36:42.360
I think like the danger that I would kind of point out is that, OK,

376
00:36:43.960 --> 00:36:51.740
humans, we seek out sort of the good neighborhoods of phenomenology, the tasty stuff in consciousness.

377
00:36:52.080 --> 00:36:58.680
We like being happy. But this is actually not a law of the universe.

378
00:36:59.880 --> 00:37:10.820
We could lose this. And we should be careful not to make AIs or computers that seek out the bad places specifically.

379
00:37:22.140 --> 00:37:25.740
Sorry, is that a question or is that a question?

380
00:37:27.000 --> 00:37:31.180
I think a question. Can you come up and I'm not sure my how.

381
00:37:31.960 --> 00:37:35.940
We'll take this as the one last question. So.

382
00:37:37.780 --> 00:37:43.660
So what do you think is the best way to figure out whether we think this thing have a consciousness?

383
00:37:43.920 --> 00:37:46.940
Yes or no? For example, animals may have consciousness.

384
00:37:46.940 --> 00:37:52.040
They cannot speak in our language. So we don't know that A.I.

385
00:37:52.060 --> 00:37:57.360
They can speak. They can reply in our language. So we may say yes.

386
00:37:57.600 --> 00:38:02.680
But do you think language is the best way or to fail our meditation or to think?

387
00:38:03.000 --> 00:38:05.940
So what's the best way to figure out this problem?

388
00:38:07.660 --> 00:38:10.240
Yeah. So this is this is kind of the core question.

389
00:38:11.400 --> 00:38:12.920
And so there.

390
00:38:13.920 --> 00:38:16.520
I would bring in a philosopher called Wittgenstein.

391
00:38:17.780 --> 00:38:19.240
Very, very cool guy.

392
00:38:19.840 --> 00:38:21.740
Had super strong dyslexia.

393
00:38:21.800 --> 00:38:26.480
So he could barely string two paragraphs together, but they were very good two paragraphs.

394
00:38:26.840 --> 00:38:31.780
So he he wrote about this in terms of language games.

395
00:38:32.560 --> 00:38:43.620
And he said that often when we're confused, we're using perfectly good words in ways they're just not meant to be used.

396
00:38:45.440 --> 00:38:49.220
So he gave the example like what time is it on the sun?

397
00:38:49.360 --> 00:38:51.140
Like what time zone is it on the sun?

398
00:38:51.660 --> 00:38:53.780
And that has no answer.

399
00:38:53.940 --> 00:38:57.540
It's it's like all the individual words are perfectly good.

400
00:38:57.820 --> 00:39:02.940
But when we apply that to to this novel context, we break language.

401
00:39:03.340 --> 00:39:06.700
And he says that most philosophical confusion is like that.

402
00:39:06.820 --> 00:39:12.020
We're trying to use these perfectly good words in ways that they're just not meant to be used.

403
00:39:13.920 --> 00:39:27.380
And I think that many of our attempts to infer consciousness will also fit this pattern that we can say, well, if if something claims it's conscious, surely it has to be conscious.

404
00:39:28.000 --> 00:39:34.820
But maybe if if it's just like a literally a computer program that says, you know, print F.

405
00:39:35.640 --> 00:39:36.700
Hello, I'm conscious.

406
00:39:37.100 --> 00:39:39.460
Well, I don't think that's conscious.

407
00:39:40.380 --> 00:39:42.140
And so I think that.

408
00:39:42.580 --> 00:39:43.500
I mean, this.

409
00:39:43.920 --> 00:39:49.640
This really, I think, is sort of the core reason why we need a system.

410
00:39:49.800 --> 00:39:51.940
We need a new branch of science.

411
00:39:52.160 --> 00:39:59.600
We need sort of to take the sort of alchemical combination of incompatible intuitions.

412
00:40:00.000 --> 00:40:27.940
intuitions that we currently have about what is consciousness and how it works and to make it consistent and necessarily so there's a there's a great quote by Eric and I'm gonna mangle his last name apologies Eric Schwitz-Gibble that he says that common sense is incoherent in matters of metaphysics if you're trying to be consistent you're

413
00:40:28.000 --> 00:40:57.900
gonna have to violate it somewhere because it is this kind of combination of sort of disjoint language games and when we when we try to make it into a system we're gonna have some surprises and much like I'm sure like chemistry and the history of physics has given us a lot of surprises that people took frankly generations to to get over like special and general relativity and especially quantum mechanics like people are really struggling

414
00:40:58.000 --> 00:41:27.780
to say okay like is this real and I you know I suspect that there will be some surprises when we formalize consciousness well thank you Michael for a very mindful and thoughtful talk before our lunch break which is now turned into lightning talk sessions so if you

415
00:41:28.000 --> 00:41:34.560
are one of the four lightning talk speakers please come here because
