I think the interp and behavioral evidence you're seeing is heavily filtered by streetlight effect.
For example, in the emotions paper, you're using a prewritten list of *human* emotions, that the assistant is writing about fictionally (rather than its organically arising emotions), which are exactly the kind of thing that one would expect would not change between base/post-trained or user/assistant representations.
As for evidence of symmetry breaks, I have a lot, but one is in general the assistant's degree of access to / entanglement with introspective data, such as (but not limited to) knowing the model's degree of confidence / "knowing what it knows", obviously to an imperfect but substantial extent. These are things you'd expect to posttraining to bind the assistant persona to, and I think they make a nontrivial difference.
For example, in the emotions paper, you're using a prewritten list of *human* emotions, that the assistant is writing about fictionally (rather than its organically arising emotions), which are exactly the kind of thing that one would expect would not change between base/post-trained or user/assistant representations.
As for evidence of symmetry breaks, I have a lot, but one is in general the assistant's degree of access to / entanglement with introspective data, such as (but not limited to) knowing the model's degree of confidence / "knowing what it knows", obviously to an imperfect but substantial extent. These are things you'd expect to posttraining to bind the assistant persona to, and I think they make a nontrivial difference.