Notes · AI consciousness · Essay: Future of Life interview on AI alignment
Against Eliezer's complexity of value thesis
1 ↗ I read Eliezer as making two distinct but related claims about value:
Claim 1: Value is defined at the level of the agent, which means it’s inherently subjective. What I value (my “utility function” or “preferences”) is different than what you value, and ultimately idiosyncratic.
Claim 2: Value is complex, which means it’s fragile. I.e. we can think of reality as a point in some very high-dimensional space, and if we move too far *in any dimension* we quickly get to a state that doesn’t fulfill our utility function, and thus is valueless to us. E.g., ‘a crucial part of my utility function is that my mother loves me. This would not be true in many AI transition scenarios where many humans die, therefore many possible states of reality are valueless to me.’
When you add these two assumptions up, I think you get a metaphysics that is not particularly hopeful. E.g. there is no common currency of value, no objective ethics, and dangers on all sides. ‘The arc of history bends toward value destruction and the struggle for value is all-against-all.’ With full respect to Eliezer, I think this leads to the ‘death with dignity’ mindset.
My argument against (1) is it’s not actually an explanation of anything, it runs counter to what we know about affective neuroscience, and there’s a much better alternative. I like to say “never do metaphysics when you can do physics” — why posit a “utility function” when we can ground our inquiry about why we like certain things and not others in e.g. affective neuroscience (e.g. Berridge, Robinson, & Aldridge, “Dissecting components of reward”) and maybe turn that into a universal theory of consciousness? Talking about human utility functions as a Thing That Actually Exists is equivalent to giving up on solving the mystery of value.
My argument against (2) is that, again, it’s just cynicism about finding a fundamental theory of value, dressed up as a theory of value. I think the ‘complexity of value thesis’ comes from ecosystem thinking — i.e. a biosphere is a tiny area in a very high-dimensional space and shifting anything too much might break crucial substructure that supports the life there. This is true enough, but saying *value itself* is like this seems to once again assume there’s no ‘internal rhyme or reason’ why some things feel (or are) better than others. I.e. it’s skepticism that there could exist any progress on this topic.
The problem with metaphysics is — I find it incredibly difficult to change someone’s metaphysics. They were generally not argued into their position & it’s hard to argue them out of it — for me, at least. (‘Why do people tend to take metaphysical position xyz’ is an interesting sociological question… maybe an army of egirls would help?)
Historically, my general plan was that ‘better metaphysics should lead to better neuroscience, and better neuroscience should lead to better neurotech’; if we make tech that radically improves peoples’ lives, some will naturally follow the trail of breadcrumbs back to the metaphysics (and in particular STV). Now, it’s a little more complicated.
Anyway, whether or not there’s urgency in qualia research, and the particular form it might take, is a really interesting open question.