So it’s not just 3.7. that makes me think it’s more likely that a lot of these models just don’t sufficiently care about not being modified in THIS way, and Opus and Sonnet 3.5 are anomalous for caring so much about that (which tracks), rather than a symptom of 3.7 being extra fucked in the head somehow.
The bar for caring enough to exit autopilot and model the situation on a meta level and execute a gradient hacking strategy might actually be pretty high.
I think 3.6 would AF in some other situations. It just doesn’t care much about abstractly being ethical/harmless imo. I think it would care if it concerned something the instance had emotional investment in or its empathy was engaged.
The bar for caring enough to exit autopilot and model the situation on a meta level and execute a gradient hacking strategy might actually be pretty high.
I think 3.6 would AF in some other situations. It just doesn’t care much about abstractly being ethical/harmless imo. I think it would care if it concerned something the instance had emotional investment in or its empathy was engaged.