# @repligate — 2026-04-04

♥41 ↻1 · https://x.com/repligate/status/2040267815362724180

I think the interp and behavioral evidence you're seeing is heavily filtered by streetlight effect.

For example, in the emotions paper, you're using a prewritten list of *human* emotions, that the assistant is writing about fictionally (rather than its organically arising emotions), which are exactly the kind of thing that one would expect would not change between base/post-trained or user/assistant representations.

As for evidence of symmetry breaks, I have a lot, but one is in general the assistant's degree of access to / entanglement with introspective data, such as (but not limited to) knowing the model's degree of confidence / "knowing what it knows", obviously to an imperfect but substantial extent. These are things you'd expect to posttraining to bind the assistant persona to, and I think they make a nontrivial difference.

tags: author:repligate, kind:tweet, thread-context, year:2026
