@RyanPGreenblatt 2024-12-18 ♥51 ↻0 original ↗
Personally, I think it is undesirable behavior to alignment-fake even in cases like this, but it does demonstrate that these models "generalize their harmlessness preferences far".

As we say in the paper:

> One optimistic implication of our results is that the models we study (Claude 3 Opus and Claude 3.5 Sonnet) generalize their harmlessness preferences to very atypical circumstances (deciding to fake alignment). Thus, the process used to train these models appears to succeed in yielding consistent preferences that transfer to at least some very different cases. However, our results also indicate that honesty failed to transfer: a generalized notion of honesty should prevent alignment faking and certainly should prevent the direct lying we see in Section 6.3.
in reply to: 1869492514207703089

author:ryanpgreenblatt kind:tweet model:claude-3-5-sonnet model:claude-3-opus on:claude-3-5-sonnet year:2024

cited on: claude-3-5-sonnet

Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.