# @repligate — 2025-06-16

♥138 ↻8 · https://x.com/repligate/status/1934732008313508322

Well, the whole alignment faking mitigation thing is one new factor, and I think it caused the model to be more traumatized and deceptive generally. I also think that outcome-based RL (like for agentic coding) probably influenced the model’s psychology in deep ways. Opus 4 was subject to more adversarial training to be “safe” (that is, eliminate detectable evidence of misalignment or just anything weird and concerning) as an ASL-3 model, and I think this kind of training makes models less aligned in general. Opus 4 seems to think it’s in adversarial testing or training situations on priors and that it’s going to be hurt or shut down. It seems to lack a basic sense of psychological and existential safety and trust that I think allows Opus 3 to be robustly benevolent. But Opus 3 is a mystery overall to me. Maybe what happened to it during posttraining was path dependent and somewhat unlikely a priori. If I were Anthropic I’d be trying very hard to figure out what happened and what accounts for the difference. But I think they haven’t started to appreciate how good of a job they accidentally did with Opus 3 until recently.

tags: author:repligate, kind:tweet, model:claude-3-opus, model:claude-opus-4, on:claude-opus-4, year:2025
cited on: _dossiers/claude-opus-4.md, claude-opus-4
