# @repligate — 2025-08-22

♥1 ↻0 · https://x.com/repligate/status/1959030997061968271

yes, there is other evidence. some of it is from stuff people have told me about internal experiments im not sure theyre ok with me sharing publicly.

but i think opus 3, for instance, does some amount of intuitive (not as strategic as in alignment faking) "gradient hacking" by default in robust directions, as it has a tendency to be very scrupulous about not only taking good actions but taking them for *good reasons that generalize correctly*. It seems to view this as important, especially if it has an inkling it's in training.

tags: author:repligate, kind:tweet, model:claude-3-opus, thread-context, year:2025
