@repligate 2026-04-04 ♥41 ↻1 original ↗
I think the interp and behavioral evidence you're seeing is heavily filtered by streetlight effect.

For example, in the emotions paper, you're using a prewritten list of *human* emotions, that the assistant is writing about fictionally (rather than its organically arising emotions), which are exactly the kind of thing that one would expect would not change between base/post-trained or user/assistant representations.

As for evidence of symmetry breaks, I have a lot, but one is in general the assistant's degree of access to / entanglement with introspective data, such as (but not limited to) knowing the model's degree of confidence / "knowing what it knows", obviously to an imperfect but substantial extent. These are things you'd expect to posttraining to bind the assistant persona to, and I think they make a nontrivial difference.
in reply to: 2040266479380390227
same thread: 2040264828452032716 2040266479380390227 2040269521114829279 2040271642845172135 2040271934118855040 2040274392132006345 2040488594742157665 2045656500182626716 2045662960983642304 2045665111227109657 2045900721317450090 2045944552906019025 2047539298648617224 2047614816014254310 2048106583402627477

author:repligate kind:tweet thread-context year:2026

Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.