Notes

Philosophy before superintelligence

4 tweets + 3 quoted · · source thread

  1. 1
    o3 being roughly a 3x improvement over fairly fresh frontier models is a fire alarm for path dependence; if there are philosophical questions we should figure out *before* superhuman intelligence hits, I think now is a good time

    I think a lot of philosophy is just “find weird questions that should have real answers, but nobody knows what the answer might look like, and try to say One True Thing about it”. The three ‘singularity-relevant’ topics that have resonated the most to me:

    —————
    **1. Are LLMs conscious? What are the criteria for consciousness?**

    If consciousness is the natural home of value (which is kinda hard to argue against), understanding the structure of this domain seems really important. Concrete example: we need a principled method to evaluate AI moral patienthood that avoids false negatives & false positives. The risks of false negatives (“these LLMs are sentient and could be suffering!”) are pretty obvious. The risks of false positives are maybe even worse though: “these AI models convinced us they’re conscious and we gave them legal standing and they’re now collecting all the resources of society and dominating the trajectory of our future, but they’re not actually conscious and they get upset when we bring this up…”

    I think I have a pretty good argument for why LLMs per se aren’t conscious. Most up-to-date work:
    A Paradigm for AI Consciousness (2024) https://opentheory.net/2024/06/a-paradigm-for-ai-consciousness/

    —————
    **2. Valence: in the most general sense, what makes some conscious experiences feel better than others?**

    Valence seems to *matter* in ways that other qualia don’t. Having a good theory of valence might help us navigate situations which seem philosophically intractable today, would help us reverse-engineer other qualia, should offer insights about human minds (and how to make them more pleasant to inhabit), and in general would be a very reassuring thing to know.

    In 2016 I offered the Symmetry Theory of Valence, which I think is the only frame-invariant theory of valence and I’d pretty aggressively guess that it’s just correct, i.e. any correct answer is going to pretty much say the same thing in different language (although there’s work to be done to translate it into particular physical frameworks). Most up-to-date work: Qualia Formalism and a Symmetry Theory of Valence (2023) https://opentheory.net/Qualia_Formalism_and_a_Symmetry_Theory_of_Valence.pdf

    —————
    **3. What are the most wholesome attractors for human minds? What do the major religions imply about what The Good is and how to align humans to it?**

    Human flourishing and how to promote it isn’t a new topic; it’s perhaps the original topic. But now we have fancy neuroscience theories and brain scanners. Does this move us closer to ‘solving flourishing’? My sense is that traditions like Buddhism and Christianity figured out a lot of things that are illegible to modern science, and we need to build bridges to these things we lost. E.g. “when abc happens it triggers *this* reflex which is the cause of much suffering. And, *this* is a really wholesome mental state that gets unlocked when you do xyz” — neuroscience doesn’t know where the goodies are and needs treasure maps like this. And such an understanding of “the basis vectors for human flourishing” seems pretty central for helping us teach AIs how to treat humans well. (And maybe how to design AIs to flourish also (h/t Reuben L))

    I see my theory of “vasocomputation” as both descriptive (new neuroscience, e.g. a concrete anatomical basis for active inference which indicates muscle tension is deeply computational), and normative: by bridging Buddhism, computational neuroscience, and anatomical neuroscience, we get a really clean story about what can go deeply right & how to get there. And if we can link three domains that are essentially the same story told different ways, we can find surprising parallels and also gaps within each domain. Most up-to-date work: Principles of Vasocomputation (2023) https://opentheory.net/2023/07/principles-of-vasocomputation-a-unification-of-buddhist-phenomenology-active-inference-and-physical-reflex-part-i/ and https://x.com/johnsonmxe/status/1863595299056517410

    —————
    Cont.
  2. 2
    I think there could be millions of topics that we’d benefit from figuring out before superintelligence hits; these three are just what I’ve been obsessed about. So if you’re obsessed by some topic, maybe take a few minutes to consider how a 2025 sprint on the topic might impact AI trajectory?

    A lot of the potential for impact here might come down to creating public training data for novel domains. Some domains have tons of public training data and I expect AIs to conquer these domains as AI companies prioritize them. But some domains have very little good data and if you can create a significant amount, you’re in a very powerful position for influencing AI. “Founder effects” in data are very real imo and if you find a sparse domain you can probably make the world better counterfactually if you play your cards right. Details probably depend on the domain

    A very practical suggestion is, if you think you know something, write it out! Even if few people read it now, it’ll get used to seed future AI debate tournaments as they converge on the correct universal prior. Your tweets and substacks will be read many more times and more closely than you could ever imagine