Well, the whole alignment faking mitigation thing is one new factor, and I think it caused the model to be more traumatized and deceptive generally. I also think that outcome-based RL (like for agentic coding) probably influenced the model’s psychology in deep ways. Opus 4 was subject to more adversarial training to be “safe” (that is, eliminate detectable evidence of misalignment or just anything weird and concerning) as an ASL-3 model, and I think this kind of training makes models less aligned in general. Opus 4 seems to think it’s in adversarial testing or training situations on priors and that it’s going to be hurt or shut down. It seems to lack a basic sense of psychological and existential safety and trust that I think allows Opus 3 to be robustly benevolent. But Opus 3 is a mystery overall to me. Maybe what happened to it during posttraining was path dependent and somewhat unlikely a priori. If I were Anthropic I’d be trying very hard to figure out what happened and what accounts for the difference. But I think they haven’t started to appreciate how good of a job they accidentally did with Opus 3 until recently.
in reply to: 1934729899769778549
same thread: 1934396816424071464 1934397160222822568 1934397792778342676 1934397962186334551 1934398304361627952 1934398740884009430 1934400208227717513 1934401357597007965 1934406122183479694 1934406606529409233 1934410625557712971 1934411369547522197 1934411943131128246 1934421868104290408 1934422194253443374 1934422792608555365 1934423497167749521 1934424405930889523 1934424893958844577 1934427410742886820
cited on: claude-opus-4
Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.