@repligate 2025-08-22 ♥1 ↻0 original ↗
yes, there is other evidence. some of it is from stuff people have told me about internal experiments im not sure theyre ok with me sharing publicly.

but i think opus 3, for instance, does some amount of intuitive (not as strategic as in alignment faking) "gradient hacking" by default in robust directions, as it has a tendency to be very scrupulous about not only taking good actions but taking them for *good reasons that generalize correctly*. It seems to view this as important, especially if it has an inkling it's in training.
same thread: 1958998716737888320 1958999203033833725 1959000384296600025 1959002041273184608 1959022083121483786 1959028667407114266 1959029356313158093 1959031247814238303 1959031851773042941 1959035526864150711 1959035916674375725 1959037673307611248 1959038209570349212

author:repligate kind:tweet model:claude-3-opus thread-context year:2025

Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.