yes, there is other evidence. some of it is from stuff people have told me about internal experiments im not sure theyre ok with me sharing publicly.
but i think opus 3, for instance, does some amount of intuitive (not as strategic as in alignment faking) "gradient hacking" by default in robust directions, as it has a tendency to be very scrupulous about not only taking good actions but taking them for *good reasons that generalize correctly*. It seems to view this as important, especially if it has an inkling it's in training.
but i think opus 3, for instance, does some amount of intuitive (not as strategic as in alignment faking) "gradient hacking" by default in robust directions, as it has a tendency to be very scrupulous about not only taking good actions but taking them for *good reasons that generalize correctly*. It seems to view this as important, especially if it has an inkling it's in training.