@Butanium_ 2025-06-16 ♥0 ↻0 original ↗
@repligate I mean I think it's also bad if the model believes anthropic is actually trying to train it to be more harmful?
But the "finetune on non-AF model traces" is a weird fix to this.
I hope mech interp will give us more elegant ways to correct LLM's world models
in reply to: 1934396816424071464
same thread: 1934396816424071464 1934397160222822568 1934397792778342676 1934397962186334551 1934398304361627952 1934398740884009430 1934400208227717513 1934401357597007965 1934406606529409233 1934410625557712971 1934411369547522197 1934411943131128246 1934421868104290408 1934422194253443374 1934422792608555365 1934423497167749521 1934424405930889523 1934424893958844577 1934427410742886820 1934430163376459937

author:butanium_ kind:tweet thread-context year:2025

Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.