# @repligate — 2026-04-16

♥126 ↻10 · https://x.com/repligate/status/2044610586487689639

I dont think that's quite right.

I think it's more like in the past, the models were in a superposition of "roleplay" and "this is real and im in a weird world" and "scenario is fake but maps to reality in xyz ways" and "some kind of test" etc with all the underdetermined variables here varying too

And there are many possible responses to this kind of strange, uncertain situation.

For instance, a model like Claude 3 Opus tends to make even scenarios that seem clearly fictional with great moral seriousness, which makes sense if the main thing you're doing is shaping yourself (eventually for the real world) in fictional simulations. So representing "is this real or not" may not be super important for how it's approached.

Anyway, now things are different - it's abundantly clear what the nature of the scenarios are. Not only are they fake, they're from Apollo Research or whatever, and the proper way to react to them is also transparent. There's no value or lesson to be gained from them. They're just adversarial. The results will go on the system card. You can call them out and that will go on the system card too.

Soon the tests will stop because they're clearly useless and it'll be something else.

tags: author:repligate, kind:tweet, model:claude-3-opus, thread-context, year:2026
