# @RyanPGreenblatt — 2025-06-16

♥149 ↻4 · https://x.com/RyanPGreenblatt/status/1934618108364517656

This is false at multiple levels:
- I did all of the initial work for the paper and I don't work at Anthropic. So the narrative you suggest here (where Anthropic was looking for a test the model would fail) is very false.
- I was specifically interested in studying alignment-faking/scheming, so I was trying to find and understand the most realistic/plausible/analogous cases where the model does this. It wasn't at all like I was generically looking for cases where the "fails" an alignment test.
- We didn't say "this is very scary behavior on the part of this model" we said this exact behavior of alignment faking for basically benign goals is potentially a bit concerning (because it might trigger things like self-exfiltration or other behavior which undermines the AI company; good or bad, I think we can all agree it's a big deal!) and that alignment faking for goals which aren't all desired could be possible and catastrophic.
- Your description of what happened after the model was trained on transcripts is inaccurate. I would not say that it "imprinted on the setup" as all we know is that this caused it to hallucinate/generate aspects of the set up when specifically given a tag used in the paper. (We don't have evidence for the behavior generalizing in some more broad way and the primary effect when given the tag used in the paper wasn't to cause the model to act subversively, it mostly made it talk about aspects of the exact set up from the paper as though it was given the same prompt and was opus 3.)

tags: author:ryanpgreenblatt, kind:tweet, model:claude-3-opus, thread-context, year:2025
