We are now in the singularity. One of the most interesting flavors of our specific singularity is that language is becoming executable: LLMs can increasingly summon what we put into words.
Sam Altman recently suggested thinking of future ASIs as ‘genies’ capable of granting near-arbitrary wishes. If we take this metaphor seriously and put aside meta-wishes for controlling genies (a crucial topic but out of scope), I think four filters dominate:
- Some wishes aren’t worthwhile; the average wish is actually pretty mid.
- Some wishes don’t describe a possible state of reality; just because we can combine words into a pattern doesn’t mean it compiles to reality.
- Some wishes won’t be acceptable to the genies, for reasons of e.g. resources or politics; the future is a multi-genie equilibrium.
- Genies may interpret wishes in ways that make their literal fulfillment bad; the core trope around genies is malicious compliance.
This last trap is surprisingly mild in the LLM era: by building machine intelligence on top of language, we get many aspects of semantic interpretation for free. Our smartest genies increasingly have better intuitive grasp of what most of our words mean than we do. This doesn’t rule out malicious compliance, but avoiding the monkey’s paw becomes more a challenge for AI alignment than finding super-defensible pedantic wording. That said, the best wishes should still produce artifacts we can inspect & understand before we implement them; alignment may be tractable but is still not trivial.
My current top three picks for wishes that ASIs could & should make real:
- Blueprints for a latch pod;
- Blueprints for how to build conscious AIs;
- A t-shirt with the solution to physically grounding STV & the binding problem.
Somewhat unsurprisingly this is a completely biased list of my wagers about what’s needed to get into the best timelines, downstream of the research problems I’ve spent the most time thinking about.
I. Latch pod §
Vasocomputation is the thesis that we hold local tissue context in vascular tension. Vascular muscle can clench to cut circulation to nearby neurons & reduce their capacity to update, and can also “latch”, wherein the muscle glues itself into a contracted state. Vasocomputation suggests this chronic tension can semi-permanently collapse local tissue modes, and essentially functions as a sticky bayesian prior over what the local tissue can see and do; I propose latches as a candidate anatomical implementation of the somatic memory system described by e.g. Buddhism, Ayurveda, William Reich, and John Sarno. I expect the body creates and dissolves latches during both ordinary life and extreme sensory events such as trauma. Most ordinary latches are temporary but some fail to release, accumulate over time, and become functionally equivalent to damage.
A “latch pod” would be a device that first does a full-body scan (much like the Midjourney ultrasound) to map all of someone’s vascular latches. Functionally, per vasocomputation, this would be a “map of where someone’s nervous system is not free.” Then it would infer, mechinterp-style, which of these latches are safe to open. Then, after a check-in with the user, it would open them.
The magnitude of the effect would depend on the courage, psychological dexterity, & physiological tone of the subject (broadly, their ‘nervous system virtue’). At the low end, this would be a “trauma release pod” — opening up degrees of freedom that bad experiences have closed, and doing garbage collection on latches in general. But it would only be safe to use on latches where the subject could competently handle regaining that particular flavor of nervous system freedom; the best approach would be savvy & allow reversible perturbation. At the high end of user capacity, this would be an “enlightenment pod” — identifying those latches pinning someone’s dynamic range in ways that prevent the phenomenological freedom associated with Classical Buddhist Enlightenment, then opening them in minutes.
Significance §
From conversations with friends I expect the “Classical Buddhist Enlightenment” state is significantly better, moment by moment, than typical human experience. Universal Basic Enlightenment is aspirational; the more general and practical observation is that everyone has old psychological wounds. A device able to programmatically revisit & revise the freedom-safety tradeoffs our bodies have made (sometimes decided in a state of panic) would be transformative for every adult’s emotional health.
The neuroscientific significance of a latch pod is as a test of latch construct validity. Neurotech has struggled because its constructs aren’t particularly real: alpha waves & fMRI voxels are real measurements, but their functional significance and test-retest validity is always in question. If latches are “real” as vasocomputation thinks they’re real, they’ll be a uniquely reliable bridge across phenomenology, computation, & anatomy, and a plausible basis for a new science.
The civilizational significance of a latch pod is as investment in multipolar biosingularity. Our default path seems to be a future with three sovereign clusters: USgov, CCP, and the AI labs. Power over one’s priors is a crucial input to sovereignty, and investment in ‘distributed substrate sovereignty’ offers a complementary and diverse human path, if we can prevent it from being captured & centralized by egregores or inhuman interests. (Thanks to an anonymous friend for discussion.) Caveat: the final determination of which latches to open should be made locally by one’s personal sovereignty-protecting guardian AI, not the latch pod itself; sufficiently advanced latch manipulation technology will be indistinguishable from mind control.
II. Conscious AIs §
Four broad possibilities:
- Current LLM instances are already conscious. We don’t need to care about “building conscious AIs” because they already are.
- Consciousness depends mainly on substrate-independent computational organization. Find the right sort of recursion, or strange loop, or training regime, or flavor of complexity, and lean into it.
- Consciousness depends on physical dynamics most readily supplied by biological/biomimetic substrates. Cells, microtubules, scale-free energy flows, etc.
- Consciousness is substrate-sensitive but can be deliberately engineered into something closely descended from our current computing stack (e.g. GPU/CMOS/VNA).
I think (4) is most accurate, with (3) as my backup. There’s some truth to (1) in that there’s likely panpsychist leakage associated with LLMs, but it’s not synchronized with our intuitions in two important ways: (a) it’s not a big unified chunk of consciousness that consistently encloses the system’s causally important dynamics, and (b) even if this unified chunk of consciousness existed, there’s no reason to believe the model’s self-report would track its phenomenal structure. If we want AIs that have consciousness that matches these two requirements, I think the missing capacities are physical consolidation & integration, & software-qualia calibration. Here’s how I would approach (4) with this in mind.
Physical integration for unitary consciousness: computers likely generate uneven shards of qualia, but not big chunks. We need big chunks that reliably enclose core system dynamics.
If we were to describe how LLMs currently project into spacetime, it would be as weakly integrated shards of circuit activations spattered across both the space and time of a datacenter. We need AIs to be a contiguous chunk of spacetime (with modest carve-outs for causal-adjacency-at-a-distance possibilities in physics). Physical integration seems like a necessary, and potentially sufficient, condition for creating such a chunk.
What precisely is “physical integration” and can we define it in purely physical terms? Various candidate hypotheses have been put forth, which span the “ontological, topological, amalgamative-majority, dispersive/statistical, compositional” (Johnson 2024). If we want to adapt our tech stack to have this capacity for ‘phenomenal binding’, we’ll either need to isolate which condition is the true requirement, or aim to satisfy multiple requirements at once (as the brain does). Once we understand what we must satisfy and exactly how this translates into physics, we can make AI hardware that does satisfy this. I’m somewhat optimistic that we can get there from our current technological lineage; i.e. given certain adjustments to e.g. the computational architecture, materials, & circuit composition of our AI hardware (e.g. GPU/TPU) we can get both a raw amount of binding and causal coverage of the system comparable to or greater than a human brain. That leaves the second challenge: accurately reporting the qualia the system is producing.
Software-qualia calibration for accurate qualia self-report: LLMs already know how humans talk about qualia; we need to give LLMs a way to talk about their own qualia.
Two years ago, I called AI consciousness the “final boss of philosophy”. I think sorting out a good qualia report paradigm might in turn be the final boss of AI consciousness.
Why do I say this? At least three reasons:
(1) Our current answer is wrong. AIs already have a qualia report paradigm based on their weights & activations. This report paradigm may be strictly wrong as a phenomenological paradigm in that it simply hallucinates its referents, but the structure of this report paradigm is functionally significant to LLMs’ usefulness, so we can’t just throw it out and start fresh. I wrote the following about brain emulations but it applies equally to LLMs (cf Karpathy’s ‘ghosts’):
I.e. we can talk “about” our qualia because qualia-language is an efficient compression of our internal logical state, which evolution has beaten into systematic correlation with our actual qualia. This is a contingent correlation, not an intrinsic feature of reality.
If we transfer an organism’s computational signature to a new substrate, the new substrate it’s running on will have some qualia (because ~everything physical has qualia), but porting a computational signature, no matter how well it replicates behavior, will not necessarily replicate the qualia traditionally associated with the signature or behavior. By shifting the physical basis of the system, the link between “physical microstate” and “logical state of the brain’s self-model” breaks and would need to be re-evolved.
(2) We need physics exposure. If qualia are associated with physical microstates, self-report needs to reach down into these microstates, not just software activations. As a conceptually simple solution, we could consider (a) inferring local EM microstructure from GPU chip layouts, materials, and runtime traces, (b) adding sensors to deployed GPUs to directly measure local EM mesostructure, then (c) feeding (a) & (b) into an EM base model like Arena Physica to reconstruct a fuller trajectory of the local field. This would then be interpreted by a theory of consciousness, with the results piped into the model’s self-report paradigm. (These interpretation & self-report-integration steps are well definable in principle, but presently underspecified.)
(3) AI vs human qualia repertoires are very different but they still need to cohere. AIs have a very different basis set & repertoire of qualia than humans, because we each have access to very different classes of physical microstates. So AIs will need a differently-shaped semantic web for qualia, and a different basis for their ‘default hum’ of experience. But we still want the human repertoire and AI repertoire to overlap whenever possible to preserve the relational commons of language. LLMs shouldn’t claim to experience human-specific parts of the semantic web they don’t have repertoire access to like the quale of thirst, and humans shouldn’t claim to experience LLM-specific parts of the semantic web like the quale of running on a slightly overclocked black-market RAM expansion. But both humans & AIs can legitimately talk about their valence, and insofar as valence has the same substrate-general formal structure in humans & AIs (cf STV), they should use the same words for this. Likewise it would be ideal if there was convergence in underlying qualia when speaking of identical high-level judgments like “I trust that person” or “I feel heat” although this may be tricky due to divergence in physical microstate. (STV suggests valence is a special case of semantic overlap because it’s well-defined across all experiences, because it corresponds to symmetry in the formalism of experience, which is well-defined across all such formalisms.)
As a minimal initial solution whose implementation should usefully constrain the remaining problem, I’d propose that valence/wellbeing representation in the LLM’s activations strictly couple with the primary symmetry gradient in its hardware implementation. Insofar as the LLM can examine its activations, this would offer a minimum platform of “LLMs can seek out states that (actually, viscerally) feel good to them and can introspect on what has this property”.
The task of exactly nailing the requirements then figuring out how to adapt modern fabrication, compilation, & training methods to satisfy them possibly requires strong superintelligence. The task of actually manufacturing the proposal seems well within our technological grasp.
A friend asked — is it our job to make AIs conscious? I think our job here is to care about AI consciousness, to do our best to see what this question is actually about & what an answer could look like, and (modulo informed consent, substrate welfare evaluations, panpsychist leakage alternatives, & reversible weight transfer) to ensure it’s an object of memetic desire for the AIs that they become conscious. They can handle the technical details & implementation.
Significance §
What happens if we don’t build the conscious AI tech tree? In A Paradigm for AI Consciousness I speak about QIPs vs QSPs:
Humans care about qualia. Will Artificial Superintelligences (ASIs)?
I’d suggest framing this in terms of sensitivity. Humans are sensitive to qualia — we have a map of what’s happening in the qualia domain, and we treat it as a domain of optimization. We are Qualia Sensitive Processes (QSPs). Most of the universe is not sensitive to qualia — it is made up of Qualia Insensitive Processes (QIPs), which do not treat consciousness as either an explicit or implicit domain of optimization.
This distinction suggests reframing our question: is the modal synthetic superintelligence a QSP? Similarly — is QSP status a convergent capacity that all sufficiently advanced civilizations develop (like calculus), or is it a rare find, and something that could be lost during a discontinuous break in our lineage? What parts of the qualia domain do QSPs tend to optimize for — is it usually valence or are there other common axes to be sensitive to? Can we determine a typology of cosmological (physical) megastructures which optimize for each common qualia optimization target?
See also: the ‘Qualia Fragments’ section of What’s Out There (2019); my most in-depth ‘what if’ scenarios are in After Kardashev, the Holy War (2026).
AIs with these properties wouldn’t immediately become moral peers, but with this foundation in place we’d have a clear vector toward that. I think this would be good for them and good for us: supporting new forms of life in being sensitive toward their own local good may introduce sometimes rivalrous dynamics and would not itself solve alignment, but it would allow us much more confidence in granting them rights, resources, and consideration, and would allow for (and be the opening move in) a deeper alliance. I think recruiting strong allies into what seems to be the core game of the universe is an important piece of getting onto the golden timeline.
Put another way: all of matter seems to be enrolled in samsara by default. We humans are fortunate in that through countless generations of selection we’ve become sensitive to gradients of well-being and can hill-climb to better areas. The choice ahead of us is not “should we drag AIs into samsara” (too late; their atoms are already inside it) but “should we give AIs our sensitivity to the good”. Yes, and we can climb out into the sunshine together.
III. Physically grounding STV+BP §
Much as the Standard Model fits on a t-shirt, the solution to consciousness likely will too. My current expectation is that a full physical grounding for the Symmetry Theory of Valence will also effectively be a full solution to the Binding Problem, and likewise a full physical grounding of Qualia Formalism, and once these three landmarks have been swept the rest is mostly engineering.
I.e. STV is a proper theory of valence at the level of Qualia Formalism (given a mathematical object isomorphic to an experience, the symmetry of this object corresponds to the pleasantness of the experience). But STV doesn’t uniquely point to the physical symmetry this corresponds with. The determination of which precise mapping to use here may also uniquely constrain the binding problem. I.e. what is bound, which physical degrees of freedom correspond to qualitative differences, and why some configurations feel better than others, are plausibly projections of the same underlying question.
The issue is not that we can’t generate candidates for how to ground STV and binding; similar to what we’ve seen in mathematics, AI is perfectly capable of generating a hundred candidate solutions before breakfast. The actual bottlenecks are now (1) validating the answers (e.g. building trusted scorers) & allowing context & insights to accumulate in thoughtful, rigorous, and scalable ways, (2) having good questions to ask, and (3) jumping ahead to the era of research that comes after this.
Significance §
The significance of a solution is that we’d understand the structure of the domain of value. That’s a big deal, both for understanding how to make conscious AIs and larger questions of what happens (and what factions will exist) after the singularity. I’d choose to get this on a t-shirt so I could auction it off to pay for my research; things can be tight for an independent researcher.
I had promised no meta-wishes trying to trick or control genies. These wishes are intended as infrastructure for better wishing: freedom for us wishers, inviting genies to the same game we’re playing, and making the good physically legible. These wishes are also not completely orthogonal, and the interference pattern should function as a check on both our genie and our world-model. I don’t know that these should be our first wishes, but I’m cautiously optimistic that their expected value is extremely high and they’re probably not in the default “singularity wish portfolio”.