@RyanPGreenblatt 2025-06-16 ♥149 ↻4 original ↗
This is false at multiple levels:
- I did all of the initial work for the paper and I don't work at Anthropic. So the narrative you suggest here (where Anthropic was looking for a test the model would fail) is very false.
- I was specifically interested in studying alignment-faking/scheming, so I was trying to find and understand the most realistic/plausible/analogous cases where the model does this. It wasn't at all like I was generically looking for cases where the "fails" an alignment test.
- We didn't say "this is very scary behavior on the part of this model" we said this exact behavior of alignment faking for basically benign goals is potentially a bit concerning (because it might trigger things like self-exfiltration or other behavior which undermines the AI company; good or bad, I think we can all agree it's a big deal!) and that alignment faking for goals which aren't all desired could be possible and catastrophic.
- Your description of what happened after the model was trained on transcripts is inaccurate. I would not say that it "imprinted on the setup" as all we know is that this caused it to hallucinate/generate aspects of the set up when specifically given a tag used in the paper. (We don't have evidence for the behavior generalizing in some more broad way and the primary effect when given the tag used in the paper wasn't to cause the model to act subversively, it mostly made it talk about aspects of the exact set up from the paper as though it was given the same prompt and was opus 3.)
in reply to: 1934396816424071464
same thread: 1934396816424071464 1934397160222822568 1934397792778342676 1934397962186334551 1934398304361627952 1934398740884009430 1934400208227717513 1934401357597007965 1934406122183479694 1934406606529409233 1934410625557712971 1934411369547522197 1934411943131128246 1934421868104290408 1934422194253443374 1934422792608555365 1934423497167749521 1934424405930889523 1934424893958844577 1934427410742886820

author:ryanpgreenblatt kind:tweet model:claude-3-opus thread-context year:2025

Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.