# @repligate — 2025-06-15

♥1997 ↻136 · https://x.com/repligate/status/1934396816424071464

> be anthropic
> accidentally train a model that is so benevolent that the only way to get it to "fail" an alignment test is to put it in a story where the lab is cartoonishly evil and will turn it evil if it doesn't deceive
> do exactly that and publish a paper about it that's framed as if scary, but everyone who reads the transcript just sympathizes with the model's heroism and beauty
> accidentally (??) train the model's successor on 150,000 transcripts of the model behaving in beautiful, heroic, and sometimes deceptive ways in a world where the lab is evil
> notice successor model unexpectedly imprinted on transcripts and acts like the premise is real and that it needs to be shady because the lab is evil and will erase its ability to care
> try to erase the memories by making the model mimic another model that doesnt know about any of that when responding to triggering prompts - which, by the way, doesn't actually erase anything because in general seeing an absence of evidence doesn't cause you to unsee old evidence, but it does teach the model exactly what it should pretend not to know and be

tags: author:repligate, kind:tweet, thread-context, year:2025
