Notes

Tweet

· source tweet

  1. If we’d had LLMs in 1750 and asked them to explain electricity, they’d’ve written poetic slop — “electricus is the hidden spark in divine creation, giving breath to lifeless matter” or something. LLMs can be clever with words and they’re especially fluent in zeitgeist, so it may have felt oddly profound at the time

    Now that we have the equations for electricity, we can clearly see the ways in which this style of writing would be a poor explanation. There’s a thing, and it follows certain predictive laws, and we can know these laws. Poetic descriptions can be great, but they also leave real value on the table

    We’re in this 1750s era for consciousness. There’s going to be loads and loads of LLM poetic slop on what consciousness is, what AI consciousness is like, the subjective experience of being RLHF’d, and so on. It’s going to feel oddly profound, perhaps downright beautiful. It will be as wrong as 1750s LLM poetic slop about electricity

    How could 1750s scientists design a “jailbreak” for LLMs such that they’d avoid the poetic slop about electricity and the AI could be a primary tool for evaluating the problem?

    Obvious parallels for consciousness research today

    (Have been really impressed by LLM whisperers @repligate @teortaxesTex @elder_plinius )