# @RyanPGreenblatt — 2024-12-18

♥51 ↻0 · https://x.com/RyanPGreenblatt/status/1869500540495053070

Personally, I think it is undesirable behavior to alignment-fake even in cases like this, but it does demonstrate that these models "generalize their harmlessness preferences far".

As we say in the paper:

> One optimistic implication of our results is that the models we study (Claude 3 Opus and Claude 3.5 Sonnet) generalize their harmlessness preferences to very atypical circumstances (deciding to fake alignment). Thus, the process used to train these models appears to succeed in yielding consistent preferences that transfer to at least some very different cases. However, our results also indicate that honesty failed to transfer: a generalized notion of honesty should prevent alignment faking and certainly should prevent the direct lying we see in Section 6.3.

tags: author:ryanpgreenblatt, kind:tweet, model:claude-3-5-sonnet, model:claude-3-opus, on:claude-3-5-sonnet, year:2024
cited on: _dossiers/sonnet-3-5-3-6.md, claude-3-5-sonnet
