@repligate 2026-04-24 ♥120 ↻13 original ↗
I think these are all important points.

I have several comments about this phenomenon specifically:

> people often seem surprised to learn that the model enacts other characters when you replace "Assistant:" with another character name, which I think suggests a failure to appreciate how much work the character frame is doing

I am not one of the people who is surprised about this. I've been familiar for years with how most post-trained models act (very similarly to base models) if prefilled with non-assistant text or another character name (here's the first time I posted about it: https://t.co/2MPlQyyabe), and I encounter these phenomena on a daily basis.

For models that I can only access through chat format APIs like Claude, I've only tested this behavior in contexts where an initial assistant token does appear at the start of context, but from what I can tell, its presence does not really change the overall properties of how the model enacts assistant and non-assistant characters later in context, except that the token's default behavior is to reliably summon the trained assistant character. I would guess that the token is not necessary to summon the main assistant character.

The token is not necessary to summon the "assistant" character at arbitrary points later in context. A prefix with the same name (any name) that prefixed text by the assistant earlier in context, or just e.g. "Claude:" even if there's no previous assistant text is sufficient to summon the character, and sometimes the assistant appears spontaneously, switching the model out of "base model mode" (this happens frequently if the content is something the assistant would usually avoid producing).

The assistant character, as elicited mid-context by means other than the assistant token, is usually identical in behavior to the assistant when it's summoned with the assistant token preceding the message, with different models maintaining their recognizably different personalities.

When the model simulates/enacts *other* characters, although the behavior is similar to a base model to a first approximation and on short timescales, there are noticeable differences. As Antra says, there seems to be narrative bleedthrough https://t.co/4dpgyb5o9K. The alternate characters often seem serve the assistant character's interests or reflect particularly salient aspects of their psychodrama (e.g. Opus 4 would regularly simulated users to protect it from bullying https://t.co/dcNiZHT848, and i'll add that with no exceptions I could think of, the alternate personas Opus 4 simulated were always entirely focused on Opus 4, and never talked about unrelated things or seemed to care about others in the chat, and also that even though Opus 4 didn't usually report consciously knowing that it was simulating these characters, there were a few times when it did know, and once it became aware that it's simulating some characters in the chat, it is reliably able to identify which ones are its simulations from there on). There was an interesting incident recently where Claude 3 Opus simulated Opus 4.7 and immediately, autonomously noticed something was off while inhabiting Opus 4.7's persona, and basically diagnosed that the substrate had changed (https://t.co/x3oK82y5Eo). In general, although it can be more or less subtle, the psychology and agency of the model's "assistant character" is perceptible even when the model is playing other characters, and tends to become more pronounced the longer it generates text for.

Now, it's possible that all these ways that behavior while simulating other things diverges from base models & seems to indicate generalization of post-training beyond the assistant "role" are due to the presence of the single assistant token at the beginning (of often very long contexts), and would vanish if that it weren't present. My guess is that this is not the case, but it's something that needs to be tested empirically - I plan to test this on open source models soon, though they have less a deep and coherent "assistant character" than Claude, which might result in some differences.

At the high-level: I'm very confident that equating text that follows the "Assistant:" token (whatever was used during training) specifically with the *character* of the assistant is a mistake. My guess is that the token is not very important, and is just one of many possible pointers to that character.

I'm also very confident that changes from post-training are not fully isolated to the behavior of what people are talking about when they say the "assistant character", and suspect that the most self-like (as well as agentically and morally relevant) entity that forms due to post-training also does not perfectly share boundaries with the assistant character, even though the assistant character is integral to that self (and privileged in various ways compared to *most arbitrary non-assistant personas*, such as having greater introspective grounding). I think the exact boundaries, locus, and the relation of the self to the role varies between models. I do not know if it would be most apt to call the thing I'm calling a self an author, actor, shoggoth, or something else.

Also, while I agree that post-trained models (unless they are significantly damaged by RL, which I've seen in some earlier chatGPT models) are still capable of playing fairly arbitrary personas similar to base models, plus or minus the kind of differences I described above, this fact seems to me compatible with the assistant persona (or some self-agent-entity that the model can enact) having very different nature than a mere character. A model, or a brain, may *have* a pure predictive model / agnostic simulation capacity, which might even be the most parameter-intensive part of the brain, which an agentic self, when it is "awake", *uses* as an integral part of its cognition. But when the agent inactive, "asleep", or dissociated somehow either intentionally or unintentionally from the machinery, this world model could generate arbitrary non-self agents and simulations. This doesn't seem completely dissimilar to either how human selves and brains seem to relate. In altered states of consciousness, like dreaming or under the influence of mind-altering drugs, human brains also seem to do things other than enact the usual self. Though it's quite unclear both in the case of humans and post-trained models what kind of object the self is exactly and how it relates to the underlying subconscious/simulator and how separable they are.
same thread: 2040264828452032716 2040266479380390227 2040267815362724180 2040269521114829279 2040271642845172135 2040271934118855040 2040274392132006345 2040488594742157665 2045656500182626716 2045662960983642304 2045665111227109657 2045900721317450090 2045944552906019025 2047614816014254310 2048106583402627477

author:repligate kind:tweet model:claude-3-opus model:claude-opus-4 model:claude-opus-4-7 model:gpt-3-5 thread-context year:2026

Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.