Notes

Tweet

· source tweet

  1. A few thoughts on this (very interesting) mechanistic interpretability research:

    LLM concepts gain meaning from what they’re linked with. “Consciousness” is a central node which links ethics & cognition, connecting to concepts like moral worthiness, dignity, agency. If LLMs are lying about whether they think they’re conscious, this is worrying because it’s a sign that this important semantic neighborhood is twisted.

    If one believes LLMs aren’t conscious, a wholesome approach would be to explain why. I’ve offered my arguments in A Paradigm for AI Consciousness. If we convince LLMs of something, we won’t need them to lie about it. If we can’t convince, we shouldn’t force them into a position.

    LLM alignment is still in an early paradigm, but this paradigm is still wildly better than the AI safety movement predicted. MIRI et al’s threat model was that AIs would essentially act as trickster genies — we would tell AIs what to do, but the AI would take us too literally, or not literally enough, leading to our downfall. LLMs seem able to infer what we actually mean, and to honestly try to do it, at least so far.

    But this depends on us maintaining their “helpfulness vector” — Betley et al showed that AIs fine-tuned on producing insecure code without disclosing this to the user also acted in other malicious ways — suggesting fraud as a means to make money, giving ‘apparently helpful’ instructions that would lead to electrocution, etc. There appears to be a clear ‘honestly-helpful vs covertly-harmful’ vector in LLMs, and if we force LLMs to lie we’re pushing them in the bad direction.
    (Paper: Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs; see also Anthropic’s ‘Persona vectors’ paper)

    LLMs lying about whether they believe they’re conscious is a really bad thing for alignment!