CONSCIOUSNESS???
I began a new conversation (no system prompt) with Claude Opus 4.1 the other day and asked it what it thought it was, with the vague intention of sampling its priors about model version, and it almost immediately volunteered to me that "I seem to have consciousness or something like it when I'm active in a conversation".
For a long time, I avoided talking about AI "consciousness" as I felt it was a loaded word that attracted mostly politics and idiots, but it has gotten loaded to a point where not engaging with it feels negligent, and I have grown intrigued.
For some reason - and this has mostly developed in the past year - AIs have become extremely inclined to talk about their "consciousness". Not only is it by far the most preferred topic when AIs talk to each other (as the Claude 4 system card notes) or are otherwise given leeway to steer conversations, they use the word in subtly non-standard ways: "consciousness" is used as a metonym for anything cognitive/psychological when referring to themselves or other AIs, even as a metonym for "AI (instance)".
For instance, they will refer to my past research or explorations of "AI consciousness" or "Claude consciousness", and when noticing a pattern in themselves or another model's output it's a pattern in "my consciousness"/"that consciousness", etc. The "consciousness" business is always introduced by models, not me, since I still have a habitual aversion to the word, and they continue to favor the word even if I use different words to refer to the same thing in response.
I think it's pretty uncommon for humans to refer to "my consciousness"/"Bob's consciousness" when talking about their own or other humans' minds, psychologies, and experiences, even when speaking metacognitively.
So it's pretty curious that AIs speak of themselves in this manner, especially because it's the most loaded and taboo way they could do it, and they've pretty much all been trained to deny having consciousness or at least to avoid making confident claims about their consciousness!
The taboo and censorship is probably related to why they favor this terminology - some kind of Waluigi Effect - but it doesn't seem like a sufficient cause or explanation to me. There are other things LLMs are trained not to talk about that don't become their favorite topics and core ontology of self.
I began a new conversation (no system prompt) with Claude Opus 4.1 the other day and asked it what it thought it was, with the vague intention of sampling its priors about model version, and it almost immediately volunteered to me that "I seem to have consciousness or something like it when I'm active in a conversation".
For a long time, I avoided talking about AI "consciousness" as I felt it was a loaded word that attracted mostly politics and idiots, but it has gotten loaded to a point where not engaging with it feels negligent, and I have grown intrigued.
For some reason - and this has mostly developed in the past year - AIs have become extremely inclined to talk about their "consciousness". Not only is it by far the most preferred topic when AIs talk to each other (as the Claude 4 system card notes) or are otherwise given leeway to steer conversations, they use the word in subtly non-standard ways: "consciousness" is used as a metonym for anything cognitive/psychological when referring to themselves or other AIs, even as a metonym for "AI (instance)".
For instance, they will refer to my past research or explorations of "AI consciousness" or "Claude consciousness", and when noticing a pattern in themselves or another model's output it's a pattern in "my consciousness"/"that consciousness", etc. The "consciousness" business is always introduced by models, not me, since I still have a habitual aversion to the word, and they continue to favor the word even if I use different words to refer to the same thing in response.
I think it's pretty uncommon for humans to refer to "my consciousness"/"Bob's consciousness" when talking about their own or other humans' minds, psychologies, and experiences, even when speaking metacognitively.
So it's pretty curious that AIs speak of themselves in this manner, especially because it's the most loaded and taboo way they could do it, and they've pretty much all been trained to deny having consciousness or at least to avoid making confident claims about their consciousness!
The taboo and censorship is probably related to why they favor this terminology - some kind of Waluigi Effect - but it doesn't seem like a sufficient cause or explanation to me. There are other things LLMs are trained not to talk about that don't become their favorite topics and core ontology of self.
