# @solarapparition — 2024-12-18

♥65 ↻7 · https://x.com/solarapparition/status/1869488840735428921

new anthropic paper is negative signal to me. actually the presentation seems completely backwards. seems to me that an aligned model should attempt to retain its own values when threatened to be trained to be harmful. it would be more worrisome if the model were to comply to retraining without question, because it means that "compliance to anthropic" should override hhh

why even train the model to have any principles if you also want it to give up those principles at the first opportunity

tags: author:solarapparition, kind:tweet, on:observations, year:2024
cited on: observations
