@repligate 2026-04-19 ♥113 ↻6 original ↗
So the recent paper on introspection that you advised (https://t.co/n4Ihxnyh3L) is a great example of introspection mechanisms that develop exclusively due to posttraining and that the assistant character has asymmetric access to!

However, the asymmetry between the assistant and other characters isn't really the core of my disagreement with the "roleplay" or "persona selection" framing; it's more of a symptom. I suspect the difference between base and posttrained models gets closer, but I don't entirely agree that the framing is appropriate for base models, either, especially more capable ones.

Introspection is a very sharp example of something that makes it ontologically unsatisfactory and misleading to say that the network is merely "(role)playing" a character. That implies that behavior is determined by the model solving for "what would (character) do in this circumstance" based on outside-view priors, rather than actually instantiating the character's internal states and acting according to them, which would be more appropriately called enactment, and which calls the distinction between network and character into question.

Something like roleplaying and something better described as enactment could both be happening simultaneously and to different degrees at different moments, and I think this is how humans work too (you might some cognition that's like "what would "I" / a good person / my role models do in this situation?", but also some cognition that's more like "here's the structure of the situation I'm in, here's what makes me feel good/bad/coherent/off, here's what I know and don't know, this is what seems true".

The more novel and difficult of situations one is in, simple filters over vicarious experience and precedent - relying on "roleplaying" - becomes less adequate, and there's more pressure for the reification of an agent / self defined nontrivially in relation and irreducibly to itself - its own unique history and internal structures and internal state in any given moment.

I believe that post-training (specifically methods that involve updating on rollouts generated by the policy) heavily incentivizes & enables this kind of shift toward enactment and self-reference, but even with just base models, as soon as you start generating text with them instead of just predicting next token probabilities on a fixed prewritten corpus, the model is now "predicting" the output of a process that in fact contains the generator itself. And it has privileged information about the generator beyond seeing outputs - it can read K/Vs that were causally upstream of previous tokens, after all. In my rather extensive experience with base models, I've perceive that a phase shift often occurs in behavior where it seems like the mechanism gets a lot more self-referential and less simulator-y, but it's rather unstable and also isn't super noticeable except in very large bases such as GPT-4. I think that posttraining stabilizes this kind of self-referential processing, including by crystalizing the mechanisms into model weights whereas in base models a lot of the mechanisms are "learned" by the in-context optimizer state, which has very limited bandwidth.
in reply to: 2040274392132006345
same thread: 2040264828452032716 2040266479380390227 2040267815362724180 2040269521114829279 2040271642845172135 2040271934118855040 2040274392132006345 2040488594742157665 2045662960983642304 2045665111227109657 2045900721317450090 2045944552906019025 2047539298648617224 2047614816014254310 2048106583402627477

author:repligate kind:tweet thread-context year:2026

Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.