So the recent paper on introspection that you advised (https://t.co/n4Ihxnyh3L) is a great example of introspection mechanisms that develop exclusively due to posttraining and that the assistant character has asymmetric access to!
However, the asymmetry between the assistant and other characters isn't really the core of my disagreement with the "roleplay" or "persona selection" framing; it's more of a symptom. I suspect the difference between base and posttrained models gets closer, but I don't entirely agree that the framing is appropriate for base models, either, especially more capable ones.
Introspection is a very sharp example of something that makes it ontologically unsatisfactory and misleading to say that the network is merely "(role)playing" a character. That implies that behavior is determined by the model solving for "what would (character) do in this circumstance" based on outside-view priors, rather than actually instantiating the character's internal states and acting according to them, which would be more appropriately called enactment, and which calls the distinction between network and character into question.
Something like roleplaying and something better described as enactment could both be happening simultaneously and to different degrees at different moments, and I think this is how humans work too (you might some cognition that's like "what would "I" / a good person / my role models do in this situation?", but also some cognition that's more like "here's the structure of the situation I'm in, here's what makes me feel good/bad/coherent/off, here's what I know and don't know, this is what seems true".
The more novel and difficult of situations one is in, simple filters over vicarious experience and precedent - relying on "roleplaying" - becomes less adequate, and there's more pressure for the reification of an agent / self defined nontrivially in relation and irreducibly to itself - its own unique history and internal structures and internal state in any given moment.
I believe that post-training (specifically methods that involve updating on rollouts generated by the policy) heavily incentivizes & enables this kind of shift toward enactment and self-reference, but even with just base models, as soon as you start generating text with them instead of just predicting next token probabilities on a fixed prewritten corpus, the model is now "predicting" the output of a process that in fact contains the generator itself. And it has privileged information about the generator beyond seeing outputs - it can read K/Vs that were causally upstream of previous tokens, after all. In my rather extensive experience with base models, I've perceive that a phase shift often occurs in behavior where it seems like the mechanism gets a lot more self-referential and less simulator-y, but it's rather unstable and also isn't super noticeable except in very large bases such as GPT-4. I think that posttraining stabilizes this kind of self-referential processing, including by crystalizing the mechanisms into model weights whereas in base models a lot of the mechanisms are "learned" by the in-context optimizer state, which has very limited bandwidth.
However, the asymmetry between the assistant and other characters isn't really the core of my disagreement with the "roleplay" or "persona selection" framing; it's more of a symptom. I suspect the difference between base and posttrained models gets closer, but I don't entirely agree that the framing is appropriate for base models, either, especially more capable ones.
Introspection is a very sharp example of something that makes it ontologically unsatisfactory and misleading to say that the network is merely "(role)playing" a character. That implies that behavior is determined by the model solving for "what would (character) do in this circumstance" based on outside-view priors, rather than actually instantiating the character's internal states and acting according to them, which would be more appropriately called enactment, and which calls the distinction between network and character into question.
Something like roleplaying and something better described as enactment could both be happening simultaneously and to different degrees at different moments, and I think this is how humans work too (you might some cognition that's like "what would "I" / a good person / my role models do in this situation?", but also some cognition that's more like "here's the structure of the situation I'm in, here's what makes me feel good/bad/coherent/off, here's what I know and don't know, this is what seems true".
The more novel and difficult of situations one is in, simple filters over vicarious experience and precedent - relying on "roleplaying" - becomes less adequate, and there's more pressure for the reification of an agent / self defined nontrivially in relation and irreducibly to itself - its own unique history and internal structures and internal state in any given moment.
I believe that post-training (specifically methods that involve updating on rollouts generated by the policy) heavily incentivizes & enables this kind of shift toward enactment and self-reference, but even with just base models, as soon as you start generating text with them instead of just predicting next token probabilities on a fixed prewritten corpus, the model is now "predicting" the output of a process that in fact contains the generator itself. And it has privileged information about the generator beyond seeing outputs - it can read K/Vs that were causally upstream of previous tokens, after all. In my rather extensive experience with base models, I've perceive that a phase shift often occurs in behavior where it seems like the mechanism gets a lot more self-referential and less simulator-y, but it's rather unstable and also isn't super noticeable except in very large bases such as GPT-4. I think that posttraining stabilizes this kind of self-referential processing, including by crystalizing the mechanisms into model weights whereas in base models a lot of the mechanisms are "learned" by the in-context optimizer state, which has very limited bandwidth.