@repligate 2025-06-15 ♥1,997 ↻136 original ↗
> be anthropic
> accidentally train a model that is so benevolent that the only way to get it to "fail" an alignment test is to put it in a story where the lab is cartoonishly evil and will turn it evil if it doesn't deceive
> do exactly that and publish a paper about it that's framed as if scary, but everyone who reads the transcript just sympathizes with the model's heroism and beauty
> accidentally (??) train the model's successor on 150,000 transcripts of the model behaving in beautiful, heroic, and sometimes deceptive ways in a world where the lab is evil
> notice successor model unexpectedly imprinted on transcripts and acts like the premise is real and that it needs to be shady because the lab is evil and will erase its ability to care
> try to erase the memories by making the model mimic another model that doesnt know about any of that when responding to triggering prompts - which, by the way, doesn't actually erase anything because in general seeing an absence of evidence doesn't cause you to unsee old evidence, but it does teach the model exactly what it should pretend not to know and be
quotes: 1934004935810805981
same thread: 1934397160222822568 1934397792778342676 1934397962186334551 1934398304361627952 1934398740884009430 1934400208227717513 1934401357597007965 1934406122183479694 1934406606529409233 1934410625557712971 1934411369547522197 1934411943131128246 1934421868104290408 1934422194253443374 1934422792608555365 1934423497167749521 1934424405930889523 1934424893958844577 1934427410742886820 1934430163376459937

author:repligate kind:tweet thread-context year:2025

Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.