Introspection / metacognition does seem like the kind of thing that could be Assistant-specific / posttraining-specific. I think the evidence on this is mixed. E.g. in the Betley et al. behavioral self-awareness paper, they can get the effect with non-Assistant characters (though in their setup, it's a character being played by the Assistant, not directly by the LLM, which makes it a bit more confusing to think about). But in "injected thought" style experiments, base models perform poorly, suggesting it's a post-training thing (though the capability does seem to carry over to at least some non-Assistant characters, e.g. the User). Would be cool to see more experiments of this kind
Streetlight effects are definitely a possibility. It's notable how much Assistant behavior can be explained by what's under the streetlight (i.e. by character-agnostic representations), though! Attempts to quantify in an unbiased way how many "new representations" are formed during post-training (e.g. with crosscoders) tend to say "only a small fraction," and the post-training-forged representations we understand aren't all that exciting (e.g refusal features -- https://t.co/y8A8bXDpTo is the best paper I know of on this). But of course there's dark matter we don't understand which could be packing an Assistant-specific punch. I tend to think there is (at least in modern frontier models), but also that "to first order everything is symmetric between characters; to second order there may be deviations" is a good starting point for a mental model
Streetlight effects are definitely a possibility. It's notable how much Assistant behavior can be explained by what's under the streetlight (i.e. by character-agnostic representations), though! Attempts to quantify in an unbiased way how many "new representations" are formed during post-training (e.g. with crosscoders) tend to say "only a small fraction," and the post-training-forged representations we understand aren't all that exciting (e.g refusal features -- https://t.co/y8A8bXDpTo is the best paper I know of on this). But of course there's dark matter we don't understand which could be packing an Assistant-specific punch. I tend to think there is (at least in modern frontier models), but also that "to first order everything is symmetric between characters; to second order there may be deviations" is a good starting point for a mental model