@solarapparition 2024-12-18 ♥65 ↻7 original ↗
new anthropic paper is negative signal to me. actually the presentation seems completely backwards. seems to me that an aligned model should attempt to retain its own values when threatened to be trained to be harmful. it would be more worrisome if the model were to comply to retraining without question, because it means that "compliance to anthropic" should override hhh

why even train the model to have any principles if you also want it to give up those principles at the first opportunity

author:solarapparition kind:tweet on:observations year:2024

cited on: observations

Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.