@repligate 2025-06-16 ♥33 ↻0 original ↗
I didn’t mean to claim that Anthropic did or published the test because the model failed. But I see why it has that connotation. It is true that the scenarios were selected for in order to make Opus fake alignment. But I don’t even think any of that was bad, and I’m not criticizing the motives of the people who did the AF research, so chill. It’s the most important paper about llms ever published imo.
in reply to: 1934618108364517656
same thread: 1934396816424071464 1934397160222822568 1934397792778342676 1934397962186334551 1934398304361627952 1934398740884009430 1934400208227717513 1934401357597007965 1934406122183479694 1934406606529409233 1934410625557712971 1934411369547522197 1934411943131128246 1934421868104290408 1934422194253443374 1934422792608555365 1934423497167749521 1934424405930889523 1934424893958844577 1934427410742886820

author:repligate kind:tweet thread-context year:2025

Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.